Headlines

Meet Zeynep Demirbas, the New York eighth-grader who tested whether AI can recognise stress; a basic machine-learning model beat ChatGPT-4o

Meet Zeynep Demirbas, the New York eighth-grader who tested whether AI can recognise stress; a basic machine-learning model beat ChatGPT-4o


Meet Zeynep Demirbas, the New York eighth-grader who tested whether AI can recognise stress; a basic machine-learning model beat ChatGPT-4o
Zeynep’s work earned her a finalist place in the 2025 Thermo Fisher Scientific Junior Innovators Challenge

Fourteen-year-old Zeynep Demirbas has found that ChatGPT-4o is less accurate than a mental health-focused AI model and a simpler machine-learning system at detecting stress in human written text.Zeynep, an eighth-grade student at Transit Middle School in East Amherst, New York, tested four different models using more than 3,500 Reddit posts. According to Society for Science, the posts had already been labelled by humans as showing stress or no stress.Her project, titled “Evaluating the reliability of Large Language Models for stress detection”, has earned her a place among the finalists in the 2025 Thermo Fisher Scientific Junior Innovators Challenge. The competition recognises young students working on science-based projects.Zeynep became interested in the project after speaking with a family friend who is a psychologist. The psychologist told her that some health insurance companies were exploring large language models (LLMs) as cheaper, 24/7 alternatives to human therapists. Zeynep wondered whether AI systems could actually be trusted to identify stress.

Testing AI models

To test the models, Zeynep used a dataset called Dreaddit. It contains 3,553 Reddit posts that human raters had labelled according to whether they contained signs of stress.She gave the data to four different models including Bidirectional Encoder Representations from Transformers (BERT), MentalBERT, Random Forest and ChatGPT-4o. MentalBERT is a version of BERT designed for mental health-related language, while Random Forest is a basic machine-learning technique that uses multiple decision trees to make predictions.Zeynep asked each model to identify which posts showed stress. She then used a measure called an F1-score to compare their performance. The score considers both how accurately a model identifies stress and how often it misses stress or wrongly labels a post as showing stress.MentalBERT performed the best in her testing, with a score of about 82 percent. BERT followed with about 79 percent. ChatGPT-4o scored about 74 percent. It also performed worse than the Random Forest model, which was included as a simpler baseline for comparison.The result surprised Zeynep because Random Forest is a much simpler machine-learning method and does not understand language and context in the same way an LLM does. “ChatGPT performing badly was ‘really surprising,’” Zeynep said.She found it particularly interesting that a simpler model could outperform an LLM with millions of parameters. “Random-forest is ‘supposed to be a very simple and old technique. So I just put it in as a baseline,’” Zeynep said, as quoted by Science News Explores. “That was very interesting; how something so small and simple was able to beat an LLM like ChatGPT that used millions of parameters and had so much coding go into it,” she added.

What results mean

Zeynep’s findings made her question whether general-purpose LLMs are reliable enough to be used for mental health assessment. A large language model is a type of machine-learning system trained on very large amounts of text. It learns patterns in language and uses them to produce responses.Her results led her to conclude that LLMs should not replace human therapists.“My project shows that LLMs are currently unreliable and unsafe to deploy as diagnostic tools,” Zeynep said.She added that the findings did not mean that LLMs are bad or cannot be useful. But, they show that general-purpose AI systems may not be suitable for a task as sensitive as assessing mental health.“We should be mindful with AI, because it doesn’t really have an acceptable grade in mental health,” Zeynep said. “That doesn’t mean that LLMs are bad, because they’re for general use. They’re not necessarily meant for mental health,” she added.Zeynep also suggested that LLMs could potentially have a different role. Instead of replacing mental health professionals, they might help identify people who are struggling and refer them to a mental health professional.

Zeynep wants to study AI bias

The project also made Zeynep interested in whether LLMs might show different results depending on a person’s gender. “One way I feel I could expand it is seeing whether LLMs carry biases toward different genders,” Zeynep said.She said she had read about cases where doctors dismiss symptoms reported by female patients because they believe women are exaggerating. She sees this as an example of personal bias and wants to know whether AI systems could show similar patterns.Since LLMs are trained using large amounts of text created by people, Zeynep believes they can also pick up human biases.Zeynep’s work earned her a finalist place in the 2025 Thermo Fisher Scientific Junior Innovators Challenge. She hopes to become a computer scientist.She said she enjoys programming but is particularly interested in how computer science can be applied to real-world problems.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *