An effort to make text-based AI less racist and horrific

[ad_1]
In another test, Xudong Shen, a doctoral student at the National University of Singapore, assessed language models based on how much they stereotype people by gender or identify them as queer, transgender or binary-free. He found that larger AI programs tended to have more stereotyping. Shen believes that those responsible for major language models should correct these errors. OpenAI researchers have also seen that language patterns are becoming increasingly toxic; they say they don’t understand why that is.
The text created by the great language models is approaching the language that has the original appearance or sound of man, but he still does not understand the things that require reasoning that almost everyone understands. In other words, as some researchers have said, this AI is a wonderful bullfight that is able to convince both AI researchers and other people to understand the words that the machine creates.
Alison Gopnik, a professor of psychology at UC Berkeley, examines how children and young people learn to apply this understanding to computers. Children, he said, are the best students, and the way children learn language comes largely from their knowledge and interaction with the world around them. On the other hand, large language models have nothing to do with the world, as their outputs are less based on reality.
“The definition of bullfighting is that you talk a lot and it seems believable, but it doesn’t make sense,” Gopnik says.
Yejin Choi, an associate professor at the University of Washington and leader of a team studying the sense of AI for the Allen Institute, has put GPT-3 through dozens of tests and experiments to document how it can make mistakes. It is sometimes repeated. Other times devils even starting with an offensive or harmful text to create toxic language.
To teach AI more about the world, Choi and a team of researchers created PIGLeT, a bad idea to understand things about the physical experience people learn to grow up in a simulated AI environment, such as touching a hot stove. This training led to a relatively small language model to outperform others in common sense reasoning tasks. These results, he said, prove that scale is not the only recipe that wins and that researchers need to consider other ways to train models. His goal: “Can we really build a machine learning algorithm to learn abstract knowledge about how the world works?”
Choi is also working on ways to reduce the toxicity of language models. He and his colleagues introduced him earlier this month an algorithm that it learns from offensive text, similar to the approach taken by Facebook AI Research; they say it reduces toxicity better than many existing techniques. Large language patterns can be toxic to humans, he says. “That’s the language out there.”
In fact, some researchers have found that trying to adjust and remove patterns bias can harm excluded people. On a piece of paper published in April, Researchers at UC Berkeley and the University of Washington found that blacks, Muslims and people who identify as LGBT are particularly disadvantaged.
The authors argue that the problem stems in part from humans labeling data that incorrectly examines whether or not language is toxic. This leads to a tendency against people who use languages differently than whites. According to the authors of this article, this can lead to self-stigmatization and psychological damage, as well as forcing people to make code changes. OpenAI researchers did not address this issue in the final paper.
A similar conclusion was reached by Jesse Dodge, a researcher at the Allen Institute’s AI Institute. He explored efforts to reduce negative stereotypes of homosexuals and lesbians by removing all texts containing the words “gay” or “lesbian” from training data in a large language model. He has seen that these language filtering efforts can lead to data sets that effectively erase people with these identities, making language models incapable of handling the text written about these groups of people or about them.
Dodge says the best way to address bias and inequality is to improve the data used to train language models, rather than trying to eliminate bias afterwards. It is recommended to better document the source of the training data and to respect the limitations of the text taken from the web, as it may over-represent people who may have access to the Internet and have time to make a website or post comments. It also requires documenting how the content is filtered and avoiding the general use of block lists to filter out-of-network content.
[ad_2]
Source link

