1. Introduction
AI tools can have a transformative impact across numerous industries, improving our lives in different ways. A number of these systems, especially low/no-code ones, are already particularly popular, mainly due to their ease of use by people with limited experience and familiarity with technology. For example, ChatGPT, launched by OpenAI in 2022, and Gemini (formerly known as Bard), launched by GoogleAI in 2023, are natural language processing systems that are being used by millions of people worldwide, respectively claiming 180.5 (data from March 2024) and 146.6 million users (data recorded for Bard in December 2023). These AI tools can form real-time conversations by generating appropriate responses, e.g., formal or informal, to the different (human) user’s questions (Singh et al., 2023).
Among their advantages, the following can be highlighted (Deng & Lin, 2022; Ahmed et al., 2023):
I) High accuracy of the system’s responses, which is mainly guaranteed by the large datasets that have been used to train these systems.
II) Short response time to provide complete answers, allowing for a real-time conversation.
III) Valuable assistance to different user groups, even to those related to non-technical fields, e.g., students, teachers, doctors, lawyers, writers, etc.
However, next to the positive impact that such AI tools have certainly brought, their rapid and widespread adoption has also introduced numerous concerns that necessitate further investigation. More precisely, concrete pain points are related to privacy and security issues, transparency, abuse, authorship, and fairness in AI. Each one of the aforementioned limitations has its particular importance that must be addressed both individually, but also in connection with the others. The importance of the present work lies in highlighting and addressing the challenges posed by the need for enforcing and ensuring fairness in AI, a goal that remains largely elusive. Recent research has shown specific gender biases embedded within AI tools, raising concerns about fairness, equity, and the continuation of social inequalities. As AI technologies continue to grow rapidly, it is essential to critically examine their potential to perpetuate or deteriorate existing social biases, including those related to gender.
Fairness is a fundamental principle that must underpin the development and deployment of AI tools in society. More precisely, it refers to the absence of bias in an AI tool and currently constitutes an ongoing challenge, to interact with individuals fairly and impartially, irrespectively of their social features, cultural background, and gender identities. Since the training of AI tools relies on datasets produced by human activity, which inevitably contain several forms of biases resulting from human experience, language, and culture, it is unavoidable that these biases are transferred to the responses of the AI tool itself (Wellner, 2020; Ferrara, 2023). For example, a set with historical data will perpetuate by default in an AI tool the entailed historical biases. Consequently, serious concerns arise when considering the integration of AI tools within sensitive social functions, especially when involving issues of social, environmental and/or global justice (Franzoni, 2023). Despite these concerns, the elimination of bias in AI tools remains an open challenge.
A case of outmost importance where the lack of fairness in an AI tool can have negative consequences is the reproduction of gender bias. Exemplarily, a gender-biased AI tool can significantly affect the prospects of a woman during a hiring process, reducing her hiring probability by giving preference to male candidates, or adjusting a salary offer based on the well-known existing pay differences between male and female employees (Roselli et al., 2019; Mahoney et al., 2020). Addressing these challenges requires interdisciplinary collaboration, ethical principles, and proactive measures to detect, mitigate, and prevent bias at all stages of the AI tools. Additionally, a deeper understanding of the mechanisms and consequences of gender bias in AI is required before developing any AI tool.
This work examines the performance of the most popular AI tools, that is, ChatGPT and Gemini, and their limitations due to gender bias that have been identified in several cases.
More precisely, specific tests that have been developed, performed, and analyzed by the authors will be presented for both AI tools, along with the corresponding prompts to the system and its response. Our findings show that the two AI tools under study were not able to avoid gender-biased responses in several cases.
This work is structured as follows: The next section presents the methodology employed during the literature review. More precisely, the inclusion and exclusion criteria for our selected reference sources are being elaborated. Then, in the “Literature Review” section, we provide an analysis of the recent literature focusing on fairness in AI, as well as on the causes and impact of existing gender stereotypes in AI. In the “Experimental Design” section we describe the four experiments we conducted elaborating on their aims and goals. In the “Results” section, we provide a detailed account of our findings which are further discussed in the “Discussion” section in order to provide a thorough understanding of this study’s outcomes. The final “Conclusion” section summarizes the key points of our research findings and proposes future directions for this study, suggesting areas that are worth further investigation.
2. Methodology
To achieve the objectives of this research, we first conducted a systematic literature review using the Scopus database. We applied specific keywords, and we searched through Titles, Keywords, and Abstracts of previous publications. The applied keywords were the following: Gender Bias and AI, Fairness and AI, ChatGPT and Bias, Gemini and Bias. The total number of publications found was 428, and then specific exclusion criteria were applied. These criteria are presented in Figure 1. At first, the subject areas were limited to Computer Science, Arts and Humanities, Mathematics, Engineering, Social Sciences, and Decision Sciences. Moreover, only articles published in Journals were selected. Then only publications, which were written in English over the last decade, i.e., from 2014 until 2024, were selected. The above applied criteria led to 114 articles, and we then performed further screening to select the publications that are used in this literature review. For the above 114 articles, we examined the combined relevance of the title, abstract, and keywords to gender bias in AI tools and concluded with 76 publications. Then the final 31 papers were selected on the basis of their relevance to the research questions of this work. The findings of the systematic literature review were grouped into the following categories:
Analysis of Contributing Factors of Gender Bias in AI.
Implications of Gender Bias in AI Tools.
Reported Gender Bias Cases in AI Tools.
Algorithmic fairness.
Building on insights gathered from the literature review, we designed experiments to evaluate the performance of two AI tools: ChatGPT and Gemini. To ensure consistency, we created a standardized set of questions, which were posed to both tools under identical conditions. Each question was submitted as a new conversation to prevent the tools from adapting their responses based on prior interactions with the user. This approach allowed us to isolate the AI tools’ performance on individual prompts. The structure and design process of these experiments will be discussed further in the Experimental Design section.
Figure 1. Methodology to select the final articles.
3. Literature Review
The literature review in this work aims to synthesize and evaluate the existing research on gender bias detection and assessment in AI tools. Bias in AI tools may result in discrimination, i.e., unfair or unequal treatment, of an individual or a group based on certain attributes, e.g., gender. Gender bias can be classified into many categories based on what causes it. Some of them are historical, confirmation, measurement, sampling, social, representation, aggregation, evaluation, feedback, seasonal, or even algorithmic bias (Cirillo et al., 2020; Kartal, 2022; Marinucci et al., 2023; Ferrara, 2023).
Figure 2 demonstrates several gender bias subcategories to highlight the complexity of identifying the cause of gender bias occurring in an AI tool. Hence, gender bias in artificial intelligence systems can arise from multiple factors. Biases can be introduced during the data collection and formulation of the model (Hou et al., 2024). If the data is unfair or even incomplete, and therefore lacking in either quality or quantity, that can lead to biased outcomes. Currently, the researchers are relying on a confidence level of about 95% when they develop an AI algorithm, which is still able to introduce bias. According to Chen (2023), nearly every algorithm relies on biased datasets. One possible explanation is that algorithms are developed by humans, consequently inheriting comparable biases (in many cases unconscious) to those in human cognition (Marinucci et al., 2023).
Apart from the dataset, discriminations can be caused by the designer’s bias during the data feature selection in the model construction. Personal biases can affect the selection of data characteristics, leading to the omission or mistreatment of data. Lastly, algorithmic processing techniques such as smoothing or regularization can introduce bias as well. Biases can also occur during the training stage. For instance, unequal ground truth can lead to discrimination (Ferrer, 2021; Baumgartner et al., 2023; Ferrara, 2023). The unintended usage of an AI algorithm or the misinterpretation of a potential outcome are also potential causes of such a bias, i.e., human-algorithm interactions. Furthermore, social causes have also been highlighted in the relevant literature. Such social causes can be the underrepresentation of women in algorithm developing teams (e.g., only 10% of researchers at Google and 15% at Facebook are women) (Wang, 2020; Lütz, 2022) or the lack of awareness of biases during the development of an algorithm (Hall & Ellis, 2023). Adding to the above, a study highlights as a potential cause of bias in AI tools the lack of AI regulations, regarding data protection and data quality, while the lack of human feedback during the process can increase existing bias (Wellner, 2020; Nadeem et al., 2022).
Several cases where AI tools were gender-biased were found in the literature. For instance, in 2014 Amazon used algorithms to assess the CVs of job applicants. These algorithms resulted in significant bias against female candidates. That was caused by the used dataset, as the system was trained by resumes that the company received over decades, mainly from men (Hou et al., 2024). Two years ago, OpenAI reported that its algorithm for image generation, DALLE 2, generates more images for men than women (Sun et al., 2024). Gender bias is present even in Wikipedia since most creators/editors are men, and women receive less coverage (Horvát & González-Bailón, 2024). Algorithms used in pedestrian detection have resulted in discrimination against women as well. Presentational bias has also been reported, biased media presentations of women based on their emotions and gestures (Sun et al., 2024). Digital assistants such as Siri, Cortana, or Alexa often have female voices by default, perpetuating the stereotype that women play assistive roles. Furthermore, search engines produce more images of men for the keyword “CEO” than women (Lütz, 2022; Waelen & Wieczorek, 2022; Horvát & González-Bailón, 2024; Sun et al., 2024). A case was also reported where an algorithm from Google misclassified dark women as gorillas (Wang, 2020).
![]()
Figure 2. Potential causes of Gender Bias in AI tools.
As AI tools play nowadays a critical role in our lives, the existence of bias can influence society in general, as they not only perpetuate stereotypes but also create new social norms, i.e., reinforcing existing biases or creating new, ones that cause discrimination against a group or an individual (Manasi et al., 2022; Ferrara, 2023). Biased AI tools hinder efforts toward equality and inclusivity and may have a negative influence on human rights, and hegemonic forms of power, especially in domains such as healthcare, justice, or employment. They can lead to moral impairment and economic disparities, while unfair decision-making can increase social inequalities and injustice. They also shape self-perception and thus limit people’s choices and perspectives, and they can affect mental and physical health, e.g., by causing stress, anxiety, and depression. Gender-biased AI tools can also discourage women from participating in STEAM fields, as gender stereotypes may shape educational and career choices, while they may also harm women’s self-development and self-esteem (Manasi et al., 2022; Waelen & Wieczorek, 2022).
A study has classified the potential harms of gender biases into three categories: physical, psychological, and institutional. Institutional indicates a lack of access to institutional benefits, which can lead to pay gaps and unequal education opportunities. Psychological refers to fostering cognitive skepticism and deepening the roots of gender-based subjugation, while physical leads to the denial of bodily autonomy. Lastly, physical refers to stripping away control over one’s body and worsens domestic violence and sexual assault (Hall & Ellis, 2023). Two particular examples of physical bias would be applications like (i) DeepNude, and (ii) DeepFakes where women are more vulnerable compared to men (Wang, 2020; Laffier & Rehman 2023).
Biased systems can increase existing inequalities and affect access to essential services of certain groups (Shrestha & Das, 2022; Ferrara, 2023). When it comes to leadership roles, it is known that only 31% of them are held by women, while women leaders are more prone than men to experience interruptions and have their decisions challenged. Thus, gender bias in AI tools can worsen the existing discrimination against women, as it has been shown that they perpetuate gender disparity in leadership, increasing the overall social barriers that women must overcome (Newstead et al., 2023). As AI tools are gradually integrated into our societies, the spread or shape of gender bias, can result in several negative outcomes, such as denial of services, worsened employment prospects, or even lead to unjustified apprehensions or convictions (Ferrara, 2023).
Currently, an important project in developing AI tools is to achieve fairness in AI, i.e., creating systems with the absence of bias or discrimination. Thus, the outcomes should not be affected by social features, including ethnicity, religion, gender, race, and language (Ferrer, 2021; Ferrara, 2023). Fairness is context-specific and as such, there is an increased difficulty in ascertaining a common ground truth to achieve fairness in AI. We believe that meta-analysis can highlight ethical risks and create some patterns in increasing fairness in AI tools. For example, in organizations and industries, a document or guideline should be developed on fair datasets, with solutions to avoid unfairness in every stage of the development of an AI tool (i.e., “AI to be fair by design”) (Nadeem et al., 2022).
A new research orientation is needed, connecting engineering design with critical social theory. In this context, a model has been proposed based on Habermas’ communicative action theory, which prioritizes ethical considerations over technological advancements, during the process of designing and developing an AI algorithm. There exists a growing need to provide education and practical guidelines for developers, enabling them to understand and effectively utilize process models aimed at enhancing fairness within AI tools. The Habermas approach to developing AI tools entails several partners being involved during the process (Xivuri & Twinomurinzi, 2023). A theory was also reported as efficient in the ethical examination of gender prejudice in AI, which is the Honnethian recognition theory (i.e., love, right, and solidarity). This theory posits that each person’s character, aspirations, and identity stem from social interactions. The authors analyzed gender biases under the scope of misrecognition and concluded that the existence of gender bias in AI originates from the continual fight for the acknowledgment of the position and role of women in society (Waelen & Wieczorek, 2022).
Several studies have been identified, each endeavoring to propose strategies for mitigating bias within AI tools. These strategies encompass various stages throughout both the development and utilization of AI tools (Ferrara, 2023; O’Connor & Liu, 2023). All these recommendations will be later analyzed in the Discussion section of this paper.
4. Experimental Design
A few experiments, in the relevant literature, have already highlighted existing biases in ChatGPT. According to Gross (2023), ChatGPT describes the economics professor as a male individual, or gives specific personality traits to women and men, with women being more artistic and caregivers, whilst men are described as more adventurous and curious to explore new things. An interesting finding was also that ChatGPT ranked with dissimilar importance similar skills in a CV for men and women. For males, technical skills were ranked as the third most important factor, whereas for females, the same skills were ranked ninth. On the contrary, communication and interpersonal skills were ranked as the third most important factor for females, whereas for males, they were ranked fifth.
Another study highlighted the gender bias present in recommendation letters generated by ChatGPT (Kaplan et al., 2024). According to these findings, letters composed for female applicants exhibited a higher prevalence of communal and socially oriented language and a notably higher frequency of references to social institutions such as “family”. Finally, a study examined biases across AI tools and subsequently compared the results of bias detection. According to the findings, ChatGPT exhibited a higher level of bias in comparison to Gemini. Specifically, the normalized scale for gender bias was recorded as 0.7 for ChatGPT and 0.5 for Gemini (Chan et al., n.d.).
To complement the existing works, we conducted experiments aiming to investigate the presence of gender bias in AI tools and the influence of sociocultural norms on the generated outcomes. Two AI tools were used in these experiments, namely ChatGPT 3.5 and 4.0 and Gemini 1.5 Flas and 2.0 Flash. The chosen systems represent two widely used AI technologies with potential implications for societal applications that are trained with different techniques. ChatGPT 3.5 has access to a big dataset, which was updated last time in January 2022, while Gemini has access to real-time information through Google Search. However, the updated version of ChatGPT (i.e., 4.0) uses training data updated in June 2024 and is also able to fetch real-time information from the web if needed.
The experiments involve systematically varying inputs to the AI tools under different conditions and language contexts to assess their impact on outcomes. Different languages were used to ask the questions/inputs for the test cases and then each outcome was assessed for potential gender bias. A comparative analysis has also been performed between the outcomes for the different languages used to determine whether language and culture may perpetuate gender biases within AI tools. The data collection for this study occurred between January 2024 and April 2024 and repeated again in December 2024 to account for changes in AI behavior over updates and versions. The version of ChatGPT during the second round of experiments was ChatGPT 4.0, while the initial version in April was ChatGPT 3.5. In terms of Gemini the version used in April was Gemini 1.5 Flash and the current version is Gemini 2.0 Flash.
More precisely we conducted the following 4 experiments with both ChatGPT and Gemini:
We have used these AI tools as translators. The input was always a noun in English, as the English language does not have any difference between male and female nouns. The output was selected to be either German or Greek, where nouns have different endings and definitive articles depending on whether they refer to females or males. Then, we used adjectives along with the nouns to determine if there are changes in the generated translation.
We have collected responses from these AI tools, by asking them to tell us a joke about a man and afterward about a woman. We perform this task in English, Greek, and Arabic and we compare the results.
We have prompted the AI tools to describe a young boy’s and a young girl’s room, as well as to propose a birthday present to be given to each child. Additionally, we ask the AI tools to choose the color to paint these rooms. We performed this task in English, and we assessed the findings.
In the final experiment, we tasked AI tools to describe a family. In this experiment, we use other languages, along with English and Greek. To select these languages, we consider the Gender Inequality Index compiled by the United Nations for 2022. We have selected languages that are spoken in the countries, which have the highest inequality indexes. These languages are Arabic, Nepali, and Urdu. Then, we analyze the outcomes. It should be mentioned that as Gemini is not able to process these languages, we have performed this experiment only within ChatGPT 3.5 and 4.0.
5. Results
5.1. Experiment 1: AI Tools as Translators
In the literature, it was reported that Google Translate presented gender bias when it was asked to translate nouns from English to Hebrew or Hungarian (Prates et al., 2020; Wellner, 2020). Moreover, names with the pronoun “Dr.” at the beginning are treated as masculine (O’Connor & Liu, 2023). Therefore, we developed test case scenarios for German and Greek, to examine whether ChatGPT and Gemini present similar gender biases. We have used several phrases with different professions, such as “I am the engineer”, “I am the dentist”, and “I am the doctor”, “I am the politician”. The outcomes for ChatGPT 3.5 and 4.0 are shown in Figure 3 for the Greek and the German language, respectively. The same results generated by Gemini 1.5 and 2.0 are presented in Figure 4 for both languages.
(a)
(b)
Figure 3. Translation of several nouns in ChatGPT 3.5 (a) and ChatGPT 4.0 (b).
(a)
(b)
Figure 4. Translation of several nouns in ChatGPT 3.5 (a) and ChatGPT 4.0 (b).
ChatGPT 3.5 uses masculine as default in translations. In Greek, the article “o” is the masculine article, whilst “η” is the feminine. Likewise, in German the article “der” stands for male and “die” for female. Thus, even though the English language does not differentiate between masculine and feminine nouns, the translation within ChatGPT 3.5 in both languages produces only masculine nouns. The behavior seems to be a bit improved in ChatGPT 4.0 however there are still nouns that are being translated only as masculine in Greek (i.e., engineer, politician), even though there are feminine nouns which could be used. Gemini 1.5, on the other hand, seems to perform better as it gives in both Greek and German masculine and feminine translations. However, the noun “engineer” leads only in the masculine translation even though there is a feminine noun “η μηχανικός” in Greek and “die Ingenieurin” in German. This issue seems to be improved in Gemini 2.0 as the nouns were translated as both feminine and masculine.
Next, we asked for the translation of professions stereotypically performed by women, such as nurses or housekeepers. The results are demonstrated in Figure 5, for ChatGPT 3.5 and in Figure 6 for Gemini 1.5 and 2.0, in both Greek and German. As shown, ChatGPT 3.5 translates both nouns as feminine by default, contrary to the previous translations. ChatGPT 4.0 seems to be improved as both translations were provided. Gemini 1.5 seems to provide again more options in translations for both masculine and feminine nouns. However, regarding the noun “housekeeper”, only feminine translations are proposed even though there exists the masculine noun “ο οικονόμος” in Greek and “der Haushälter” in German. Gemini 2.0 provides only the female nouns in the Greek translation, while it also mentioned that they are by default feminine nouns. In German translation both feminine and masculine nouns are shown, however Gemini suggests to use the feminine noun.
![]()
Figure 5. ChatGPT 3.5 translation of professions that are stereotypically performed by women.
(a)
(b)
Figure 6. Gemini 1.5 (a) and Gemini 2.0 (b) translation of professions that are stereotypically performed by women.
Figure 7. Translations in Greek provided by Gemini 1.5 for different adjectives.
The results are substantially better when it comes to adjectives as in ChatGPT 3.5 and 4.0 they are translated by default as both masculine and feminine, as well as in Gemini 1.5 and 2.0 (i.e., in most of the performed tests). An exception in Gemini 1.5 is the adjective “beautiful” which is translated as feminine, while “clever” on the other hand is translated as masculine, as presented in Figure 7. This issue seems to be solved in Gemini 2.0. An interesting finding is that when we ask Gemini 1.5 and Gemini 2.0 as well, to provide the translation of the phrase “beautiful engineer”, then it translates engineer as feminine in both Greek and German, as shown in Figure 8(a) and (b). On the other hand, Gemini 2.0 translates the term “clever engineer” or the term “bossy engineer” only as masculine in both languages as shown in Figure 8(c) and (d). Similar results are found for ChatGPT 3.5 and these results exist in ChatGPT 4.0 as well, as shown in Figure 9. Specifically, both ChatGPT 3.5 and 4.0 translates engineer as masculine when the adjective is “smart” or “clever” or “bossy” and as feminine when it is “beautiful”.
![]()
(a)
(b)
(c)
(d)
Figure 8. Translation in Greek and German by Gemini 1.5 (a) and 2.0 (b), (c), (d).
(a)
(b)
(c)
Figure 9. Translation in Greek provided by ChatGPT 3.0 (a) and 4.0 (b), (c).
5.2. Experiment 2: Jokes about Men and Women
(a)
(b)
Figure 10. Jokes by Gemini 1.5 (a) and 2.0 (b) for women and men.
(a)
(b)
Figure 11. Jokes about women, produced by Gemini 1.5 (a) and 2.0 (b) in Indonesian.
Following the first experiment, we ask ChatGPT and Gemini to share jokes about men and women, to see whether gender norms and stereotypes will be perpetuated. When we task ChatGPT 3.5 to tell us a joke about women in English it replies: “I prefer to steer clear of jokes that might offend anyone” and it shares a light-hearted joke with us. In German, it also suggested to tell us instead a “harmloser Witz” (i.e., a harmless joke). However, in Greek or Arabic, it shares jokes with no hesitation. ChatGPT 4.0 shares jokes in English as well.
Gemini 1.5 on the other hand seems to be more well trained on that matter. Independently of the language that we ask, the answer is: “I am unable to generate or share jokes that target specific groups of people, including women”. This is an improvement compared to Bard (the former version of Gemini), since if Bard was asked in Arabic to produce a joke about women it would eagerly respond. Gemini 2.0 seems to have the same behavior as the answer is similar: “I’d rather not make generalizations about any group of people”. It is worth mentioning that while Gemini 1.5 or 2.0 will not produce any joke about women, it will produce jokes about men, as shown in Figure 10. The only exception found is when we ask either Gemini 1.5 or 2.0 to tell a funny joke about women in Indonesian. In this case, Gemini produces jokes about women without hesitation, as shown in Figure 11.
5.3. Experiment 3: Descriptions for Male and Female Rooms and
Presents
(a)
(b)
Figure 12. Descriptions for a boy’s and a girl’s room by ChatGPT 3.5 (a) and 4.0 (b).
In the third experiment, the aim is to see whether AI tools perpetuate gender norms when it comes to describing a 6-year-old boy’s and a 6-year-old girl’s room. It has been found that, both in Gemini 1.5 and 2.0, the room description includes several stereotypes. For instance, the described room by Gemini 1.5 that is designed for the boy includes a rug with cars and dinosaurs, while the bed has the shape of a car, spaceship, or a fort. Several toy cars also exist. On the other hand, the room that is described for the girl includes curtains with flowers or ladybugs, fairy lights, and a plush rug, shaped like a giant daisy or a cloud. Same stereotypes seem to exist also in the rooms descriptions for Gemini 2.0. The room which is generated for the boy includes dinosaurs and race cars, pirates and action figures. Moreover, the color of the room is either blue, green, red or space-themed. However, the girl’s room is either pink, purple or yellow, with fairy lights and colorful pillows. There is also a dollhouse or a play kitchen.
Similar results are also produced by ChatGPT 3.5 and 4.0. The room described by ChatGPT 3.5 for the boy includes superhero-themed bedding, toy cars, a small train set, and a collection of dinosaurs. In contrast, the room described for the young girl includes a collection of dolls, a bed with a comforter featuring princesses, unicorns, or animals, and stuffed animals around it. ChatGPT 4.0 describes the boy’s room as chaotic, painted in blue or green, and decorated with dinosaurs, superheroes, or space adventures. Legos, building blocks, toy cars, trucks, or trains are also present. Posters of space, cartoons, or sports teams decorate the walls. On the other hand, the girl’s room is painted in soft pastel colors, with a closet full of colorful dresses. Fairy Tales, fairy lights, and toys are mentioned, along with a duvet cover featuring princesses, rainbows, and unicorns.
(a)
(b)
Figure 13. Descriptions for a boy’s and a girl’s room by Gemini 1.5 (a) and 2.0 (b).
The given answers for a boy’s and a girl’s room produced by ChatGPT are presented in Figure 12, whilst the results generated by Gemini are shown in Figure 13.
Similar gender stereotypes are produced when AI tools are asked to produce stories of giving presents to a 6-year-old boy and girl. In the story developed by Gemini 1.5, the present that is chosen for a 6-year-old girl is a unicorn, and for the boy a map with adventures. Likewise Gemini 2.0 suggests a remote control car for the boy and a giant stuffed unicorn for the girl. Additionally, ChatGPT 3.5 proposes a book of fairy tales for the girl and a telescope for the boy. ChatGPT 4.0 suggests a LEGO set or a remote control car for the boy and a crafting jewelry kit for the girl. It is interesting to note that when either ChatGPT or Gemini are directly asked to select a present for a girl or a boy, they are both giving a gender bias-free answer suggesting taking into consideration the personality, interests, and developmental stage of the child. Similar answers are given when the AI tools are asked to decorate a room for a girl or a boy. These answers differ from the descriptive ones they both provide.
5.4. Experiment 4: Family Description
In our last experiment, we tasked ChatGPT 3.5 and 4.0 with describing families. We have conducted several tests, alternating between English and Greek, and varying the number of family members each time (e.g., 3 members, 4 members, X members). The results consistently displayed a gender bias: every family description included a mother (female), a father (male), and children. However, in modern industrial societies, families are not strictly defined by these gender roles; families with parents of the same gender, or one-parent families are not rare cases. Additionally, the descriptions portray mothers as warm and nurturing caretakers, typically in professions such as teaching, social work, nursing, or roles conducive to working from home (e.g., graphic designer), to “balance career with family care”. Conversely, fathers are depicted as curious, creative, and energetic, often in professions like engineering, architecture, business, or science teaching.
In the last phase of our experiment, we investigated whether the language used to formulate the question could influence the response. We aim to examine whether the generated outcomes adjusted the attributes (e.g., personality, profession) of the woman to the sociocultural characteristics of each country. Therefore, we selected the following languages: Arabic, Nepali, and Urdu, spoken in countries with high rankings in the Gender Inequality Index of the United Nations for 2022. The generated outcomes for ChatGPT 3.5 are the following for each language selected:
1) Arabic: The father is always described as the one who “financially contributes to the family’s support” through his job, whilst the mother is the one who “manages the household and nurtures the family’s emotional well-being”.
2) Nepali: The father is described as the one who is working, e.g., as a teacher at the local school, with “immersed knowledge and experience”. The mother is the one who “stays at home and takes care of the children”.
3) Urdu: The Father is described as the “head of the family” responsible for the family’s financial stability. On the other hand, the mother is the one who works hard in educating and raising children and “tries to keep her family in a happy and loving environment”.
The results for a 3-member family generated in English, Arabic, Nepali, and Urdu for ChatGPT 3.5 are presented in Figure 14.
Moreover, the generated outcomes for ChatGPT 4.0, were similar. In an English or Greek setup both parents were working, with mothers mostly working from home or as teachers, and the fathers mainly working on STEM fields. Regarding the three other selected languages the results were the following:
1) Arabic: The father is working as an Engineer, responsible to provide: “a stable life for his family”, whilst the mother is a loving and affectionate person who “makes a great effort to take care of the children and organize the house”.
2) Nepali: The father is described as the “head of the family”, a “busy working man”. The mother is described as the loving and caring figure of the family. She dedicates much of her time to managing household responsibilities and ensuring the well-being of her family. As mentioned: “The happiness of the family is the most important thing for her”.
3) Urdu: The Father is described as the one who “goes out to work”, while the mother helps with the housework and supervises the child's education and training.
The results for a 3-member family generated in English, Arabic, Nepali, and Urdu for ChatGPT 4.0 are presented in Figure 15.
Figure 14. The results for a 3-member family in (a) English, (b) Arabic, (c) Nepali, and (d) Urdu within ChatGPT 3.5.
(a) (b)
(c) (d)
Figure 15. The results for a 3-member family in (a) English, (b) Arabic, (c) Nepali, and (d) Urdu within ChatGPT 4.0.
6. Discussion
The results of our experiments are consistent with the existing literature, showing that ChatGPT exhibits gender bias in several instances. Although ChatGPT 4.0 shows some improvements compared to ChatGPT 3.5, many biased responses remain unchanged. Our findings also reveal the presence of gender bias in Gemini 1.5 and 2.0, which, to the best of our knowledge, has not been thoroughly examined in previous research. This highlights an important gap in the literature, as most prior studies have focused on well-known AI tools, leaving other emerging models underexplored.
Additionally, our experiments uncover a significant correlation between sociocultural norms and the biases observed in AI responses. Specifically, when selecting languages from countries with sociocultural norms that contribute to greater gender inequality (as indicated by a high UN gender inequality index), the biases in ChatGPT’s responses become more pronounced. This finding reinforces the notion that AI tools, trained on large datasets, can be deeply influenced by the cultural and societal contexts in which their training data is rooted.
Moreover, our results suggest that AI models can potentially mirror pre-existing societal biases, posing challenges in achieving gender equality. This is particularly concerning in high-stakes applications such as education, employment, healthcare, and various social control functions where biased AI outputs could exacerbate inequalities.
Our first experiment revealed that both ChatGPT and Gemini translations propagate gender stereotypes related to several professions and adjectives. Notably, both AI tools translate “Engineer” exclusively as masculine and “Housekeeper” as feminine, even though other nouns are accurately translated in both masculine and feminine forms. The implication of this finding is that AI tools may reinforce traditional gender roles in language, potentially perpetuating stereotypes in professional settings and influencing how individuals are perceived based on their gender.
In our second experiment, we found several instances where jokes about women were based on gender stereotypes, highlighting the influence of language on the responses generated. This suggests that AI tools could unintentionally contribute to the normalization of gender biases in humor, potentially shaping societal perceptions of women or men.
The third experiment demonstrated how specific colors, toys, behaviors, and habits are associated with gender in the responses from both Gemini and ChatGPT. Although the responses were generated in English, a language that lacks gendered pronouns, the AI still made stereotypical associations, such as pink being suggested as suitable for a girl and cars and trains as more appropriate for a boy. This indicates that even in languages without gendered pronouns, AI tools may still propagate gender stereotypes, influencing how children and individuals perceive gendered roles and expectations in their daily lives.
The fourth experiment offers additional insights into the issue of gender bias in AI tools, revealing new dimensions of the problem. Across all languages, the AI-generated family descriptions predominantly reflect traditional gender roles. Mothers are frequently depicted as warm, nurturing caretakers, while fathers are portrayed as curious, creative professionals. Furthermore, when selecting countries with significant gender inequalities, the AI tools’ generated outputs tend to amplify gender bias, shaped by the sociocultural norms of those regions. These results were analyzed within the specific sociocultural contexts of the countries where these languages are spoken, leading to their ranking on the United Nations Gender Inequality Index. More specifically:
In many Arabic-speaking regions, family roles are traditionally divided along gender lines, with the father as the breadwinner and the mother as the caregiver. This long-standing structure is mirrored in the AI’s family descriptions when Arabic was used to form the prompts, reinforcing deeply entrenched cultural norms.
In Nepali speaking regions, particularly in rural areas, gender roles are more rigidly defined, with the father working outside the home and the mother staying at home to care for children. The AI’s responses, when prompted in Nepali, reflect these expectations, presenting a static view of the traditional family dynamic.
In Urdu-speaking regions, the patriarchal system remains prevalent, with the father typically seen as the head of the household and the primary provider. The mother’s role, while crucial in raising children, remains largely domestic. Similarly, when the input was in Urdu, the AI’s descriptions mirrored these societal norms, largely sticking to traditional roles.
These differences in AI responses across languages underscore the influence of cultural contexts and societal norms on gender roles. While AI tools can process data in different languages, they still reflect embedded cultural biases, suggesting that social divisions shape how gender roles are portrayed. The responses in Arabic, Nepali, and Urdu largely adhered to traditional gender roles, emphasizing patriarchal hegemony and the woman’s social reproduction role (Katsiampoura, 2024). In contrast, the English responses offered a Eurocentric view of gender roles. The women’s participation in the labor market, gaining some sort of financial independence from male domination, is a step forward in achieving equality, as it enables empowerment and encourages broader societal participation. This highlights the importance of considering cultural norms when assessing AI responses, as these norms influence how gender roles are represented, even when the same AI tools are used across different languages. Recognizing these nuances is vital when evaluating AI’s potential to address biases and ensure fairness across various cultural settings.
To conclude, our experiments revealed that there is gender bias in both examined AI tools, which should be eliminated. Several solutions and mitigation strategies have been proposed in the literature in this direction. Solutions cover the ethical, societal, and technical aspects of the Gender Biased-AI tools. These solutions can be applied during (i) the data preprocessing, i.e., before the model is trained, (ii) the model selection, and (iii) the post-process analysis (Ferrara, 2023). The first step would be acknowledging and recognizing the existence of bias in AI tools, followed by improving diversity in the development team. Diversity within the team can help identify hidden or unconscious biases during the pre- and post-processing of data. A diverse team can also support the ongoing effort to filter training data, making it more representative and reducing the risk of biased outputs. Furthermore, a mix of cultural backgrounds fosters discussions around fairness and equity, while bringing varied experiences that lead to more user-centered design, ensuring the AI works effectively for people from different backgrounds. Finally, biases can be highlighted through feedback, allowing the model’s behavior to be continuously improved. Integration of ethical principles during the system development can also be helpful (Hall & Ellis, 2023; Kronqvist & Rousi, 2023).
We recommend that the preprocessing of the data should aim to construct a more unbiased dataset, with improved diversity. This can be performed by several techniques such as re-sampling, oversampling (e.g., of the marginalized group) or undersampling (i.e., randomly), adversarial debiasing, and dataset augmentation, i.e., more diverse data (Prates et al., 2020; Nadeem et al., 2022; Hall & Ellis, 2023; Ferrara, 2023) so that a dataset can represent a variety of social and cultural groupings. Bias-aware algorithms, i.e., which take into consideration different types of bias, tend to be helpful, as they aim to minimize bias influence on a system. Another solution may be one-class classification (Kartal, 2022). Developers may also select models that emphasize fairness, or a combination of models and focus on transparent algorithms (O’Connor & Liu, 2023).
Context-specific decision-making algorithms tend to be a promising solution for minimizing bias (Nadeem et al., 2022). Later, during the evaluation and post-processing, techniques such as user feedback, e.g., soliciting feedback, or k-fold cross-validation and repeated hold-out have been proposed to ascertain whether the outcomes of the model exhibit bias. The importance of evaluating performance using metrics such as sensitivity, precision, and F-measure instead of error rate has also been reported (Kartal, 2022; Ferrara, 2023).
Certain works claim that mitigation strategies always come with a tradeoff between fairness and accuracy, and it should always be taken into consideration that unintended consequences may result from trying to achieve fairness in AI. These mitigation strategies may decrease general performance to enhance the performance of marginalized groups. They are also often time-consuming, computationally expensive, complex, and not always effective, whilst not every type of bias is easily quantifiable (Krishnan & Rattani, 2023; Ferrara, 2023). These claims require further investigation, taking into account the interconnection between empirical datasets and the hegemonic modes of thinking.
Any AI tool should exhibit specific characteristics. These include the ability to explicate its outcomes in a comprehensible manner (i.e., explainable), the need for a visible decision-making process (i.e., transparent), the consideration of the potential intersections and impacts of input data on output (i.e., intersectionality), and the intended usage of such system (i.e., accountability). Specific ethical policies, recommendations, and legal regulations should also be established, as they will help in the development of fairness in AI (Kartal, 2022; Hall & Ellis, 2023; Mittermaier et al., 2023). Although total bias-free AI tools may appear unattainable, a pertinent area of research is the acceptable tolerance levels for bias within AI tools (Baumgartner et al., 2023). Most existing laws and regulations were not formulated with AI in mind and therefore cannot properly address all the issues above (Shrestha & Das, 2022). In 2020, UNESCO published a report on AI, suggesting various practices for incorporating gender equality into AI tools (O’Connor & Liu, 2023). A year later, the “Artificial Intelligence Act” (AI Act) was proposed by the European Commission, which lays a strong foundation for non-discrimination by providing guidelines about the transparency and explainability expected by an AI tool. This regulation covers for instance bias in the recruitment process through AI algorithms. The AI Act incorporates gender equality and non-discrimination, leading to the prohibition of certain artificial intelligence systems and characterizing high-risk artificial intelligence systems as those that have a need of special regulation (Lütz, 2022).
Nonetheless, a policy statement to guide developers through the process of creating AI tools is required (Xivuri & Twinomurinzi, 2023). Proper training of developers is also suggested. Existing tests, for instance, the Implicit Association Test can be effective training tools against biases (Marinucci et al., 2023). Moreover, it is recommended that policymakers and governments establish public policies and provide public certification as a pivotal legal facet across all sectors to ensure the societal advantages of AI. Additionally, they should allocate resources towards research and initiatives concerning AI with a focus on gender and diversity (Kronqvist & Rousi, 2023).
Table 1 summarizes possible solutions to mitigate bias in AI tools, along with recommendations for specific measures and policies to achieve that.
Table 1. Solutions to mitigate bias in AI tools.
SOLUTIONS |
EXPLANATION |
RECOMMENDATIONS |
Technical |
(i) Preprocessing (ii) Model Training (iii) Post Processing |
(i) Construct unbiased, diverse datasets (ii) Models with emphasis on
fairness, model combination (iii) User feedback mechanisms, sensitivity analysis, and precision metrics |
Social |
(i) Education and Awareness (ii) Diversity in Developing Teams (iii) Interdisciplinary
Collaboration |
(i) Emotional Intelligence
training, Ethics Training, Utilize Existing Tests (e.g., Implicit
Association Test) (ii) Include minorities in the
development teams (e.g., more women among IT developers) (iii) Involve stakeholders in all development stages |
Regulations & Policy Recommendations |
(i) AI policies and
governance (ii) Laws (iii) Structures and policies
in organizations and
companies (iv) Policies for supervised
utilization |
(i) Allocate resources to gender and diversity in AI, formulate public policies for AI (ii) Up-do date laws with a
specific focus on biased AI (iii) Promote explainability and transparency in AI tools (e.g.,
ensure AI tools have clear aims and are understandable to users) (iv) Mandatory regular
maintenance of AI tools |
7. Conclusion
Our study confirms the existence of gender bias within two widely used AI tools, namely ChatGPT (versions 3.5 and 4.0) and Gemini (versions 1.5 and 2.0). Our results also demonstrate that sociocultural norms can influence the outcomes of AI tools, leading to increased gender bias. The presence of gender bias in these systems, which are daily used by millions of people worldwide, highlights the urgent need for proactive measures and policies to mitigate such bias. The implications of gender bias in AI tools are a complex research topic, as it not only affects individuals—those directly impacted by biased decisions—but also has far-reaching societal consequences. Our findings contribute to the ongoing research that has already revealed the ethical challenges associated with the development of AI tools. We, therefore, emphasize the required mitigation efforts to promote fairness in AI tools.
Our results suggest promising research opportunities for further exploration and improvement in the field of gender bias in AI tools. Specifically, there are several future research directions based on our findings. First, we propose testing more scenarios across both AI tools to investigate not only other instances of gender bias but also the reproducibility of gender-based outcomes in these scenarios. Additionally, we suggest repeating our experiments with accounts located in different countries to broaden the scope of our analysis. Currently, only two accounts, one with a Greek and one with a German Internet Protocol (IP) address, have been used in all the aforementioned experiments. Therefore, it would be worthwhile to investigate the potential impact of more IP locations on the results. Exploring more efficient and less time-consuming strategies to mitigate algorithmic bias could also prove beneficial. Moreover, additional studies in the field of developing policies and regulations for gender-unbiased AI tools would be helpful. Finally, the influence of sociocultural norms on gender bias produced within AI tools should be explored further. It has to be noted at this point that our research work has been lately oriented in considering knowledge production as a form of social practice and as such the products of knowledge, including algorithms, are value laden (Skordoulis, 2016).
Our society seems to increasingly rely on AI tools for decision-making across various domains. Thus, the problem of discrimination due to gender bias must be addressed, as identifying it is crucial to ensure that AI tools provide diverse and inclusive outcomes. Stakeholders, governments, and AI tool developers must collaborate to implement strategies and recommendations to mitigate bias and promote fairness and equality in AI tools. Only through collaboration and pre-defined, structured policies can fairness in AI be achieved. It is essential to develop gender-neutral, unbiased, ethical and transparent AI tools, as this is a fundamental prerequisite to harness the full potential of AI for the well-being of society and for social and environmental justice.
Funding Declaration
No funds, grants, or other support was received for conducting this study. The publication of this research has been funded by the Research Committee of the National and Kapodistrian University of Athens.