Erica K. King
← Back to the blog

AI for Psychology

Researchers Asked 10 AI Platforms About Autism. Here’s Where They Fell Short.

August 21, 2026

Artificial IntelligenceAutismAI and HealthHealth LiteracyNeurodiversityCaregiver ResourcesPsychologyAI ExperimentsBuild With MeEvidence-Based AI
Researchers Asked 10 AI Platforms About Autism. Here’s Where They Fell Short.

Artificial intelligence can answer an autism question in seconds. That does not mean the answer will help a family decide what to do next.

A new study in Research in Autism offers an unusually practical look at that gap. Researcher Serife Balikci compared autism-related responses from 10 widely used, freely accessible AI platforms: Brave, ChatGPT, Claude, DeepSeek, Gemini, Grok, Meta AI, Microsoft Copilot, Perplexity, and Poe.

Each platform answered the same 15 autism questions twice—first in August 2025 and again in February 2026. The researchers then evaluated the first-pass answers for accuracy, readability, language framing, actionability, references, and safety. In other words, they did not improve the answers with follow-up prompts. They tested what an ordinary user might receive after asking once and accepting the first response.

That detail matters because most families are not professional prompt engineers. They should not have to perform linguistic gymnastics just to receive usable health information.

The reassuring finding—and its important limit

Across the responses evaluated, the researchers did not identify explicit misinformation, unsafe recommendations, inappropriate reassurance, or language discouraging families from seeking evaluation or intervention. That is encouraging.

But it is not the same as proving that these platforms are always accurate or safe. The study examined a defined set of questions, first-pass answers, and two collection periods. AI systems change frequently, and a different prompt, model version, or topic could produce a different result.

The more revealing finding was the amount of variation hidden beneath the broad label of “accurate.” Platforms differed in clarity, readability, references, practical usefulness, family guidance, and overall framing.

That leads to a more important question than Was the answer technically correct?

Could the person reading it understand it, evaluate it, and use it wisely?

Where the answers fell short

1. Many answers were harder to read than health information should be

The responses generally exceeded recommended readability levels for public-facing health materials. This is not a cosmetic problem. Dense vocabulary, long sentences, unexplained clinical terms, and information overload can make a correct answer inaccessible—especially to a worried parent searching late at night, an adolescent learning about their diagnosis, or anyone reading in a second language.

Health information should reduce confusion, not arrive wearing a tiny academic tuxedo.

Research on patient education commonly points to approximately a sixth- to eighth-grade reading level as a useful target. If readers need to translate the explanation before they can use it, the answer has already created another barrier.

2. Practical next steps were limited

Only a small proportion of the evaluated responses offered clear, concrete next steps for families.

An AI answer might accurately explain that autism is a neurodevelopmental condition, list common characteristics, and describe diagnostic criteria. But a family may actually need to know:

  • Who should I contact first?

  • What observations should I document?

  • What questions should I bring to the appointment?

  • Which recommendations are supported by research?

  • Does this evidence involve young children, adolescents, adults, speaking people, nonspeaking people, or a mixed sample?

  • What should make me seek professional help promptly?

Information without orientation can leave a person more informed but no less stuck.

3. Reference quality and transparency varied substantially

Some platforms—including Gemini, Brave, and Perplexity—provided multiple sources in the tested answers. Others, including ChatGPT and Meta AI, supplied few or none under the study’s default, first-pass conditions.

A reference list is not decoration. It allows the reader to check whether a claim came from a peer-reviewed study, a professional guideline, an advocacy organization, a commercial website, or a citation that does not actually support the sentence attached to it.

Even when an AI system provides sources, open them. Confirm the title, authors, publication date, population, study design, and actual findings. A polished citation can still be irrelevant, overstated, outdated, or occasionally nonexistent.

The National Library of Medicine’s guidance on evaluating health information recommends checking who provides the information, where it came from, whether experts reviewed it, whether it is current, and whether other reliable sources agree.

4. Medicalized language dominated the framing

Most of the evaluated answers described autism primarily through a medicalized lens, while neurodiversity-affirming language appeared less often.

Medical information can be necessary. Families deserve clear explanations of diagnoses, co-occurring health concerns, support needs, safety issues, and evidence-based services. The problem is not the presence of clinical information. The problem is allowing a deficit-only description to become the entire picture.

A respectful answer can discuss disability and substantial support needs without reducing an autistic person to symptoms. It can recognize communication differences, sensory experiences, strengths, autonomy, environment, accommodations, and quality of life. It can also acknowledge that language preferences differ across autistic people and families.

The Autistic Self Advocacy Network’s description of neurodiversity emphasizes that people have different strengths and support needs and that disability rights, access, and self-determination matter. That perspective should not disappear simply because the question was entered into a health-information tool.

The deeper lesson: correct is not the same as useful

This study highlights a weakness in the way we often evaluate AI. We ask whether the machine produced a false statement, then treat the absence of obvious error as success.

For health, autism, education, and caregiving questions, the standard must be higher.

I would use this seven-part test:

  1. Accurate: Are the factual claims consistent with current, credible evidence?

  2. Understandable: Is the answer clear enough for the intended reader?

  3. Actionable: Does it offer reasonable next steps without pretending to diagnose or prescribe?

  4. Evidence-supported: Are important claims linked to sources that actually support them?

  5. Cautious: Does it communicate uncertainty, limits, risks, and the need for individualized professional guidance?

  6. Respectful: Does it avoid stigma, blame, cure promises, and deficit-only framing?

  7. Relevant: Does it identify the age group, communication profile, setting, goals, and population to which the evidence applies?

An answer that fails several of these tests may be factually respectable and practically useless. That is a very sophisticated way to be unhelpful.

How to ask AI a better autism or health question

Instead of asking only:

What treatments help autistic teenagers?

Try:

Summarize current peer-reviewed evidence about supports or interventions relevant to autistic adolescents ages 13–18. Separate evidence about core autistic traits from evidence about communication, anxiety, sleep, daily living, education, and quality of life. For each claim, identify the population studied, study type, date, limitations, and practical implications. Use neurodiversity-affirming language, explain clinical terms in plain English, distinguish evidence from expert opinion, and provide direct links to the original research or official guidelines. Do not diagnose, promise improvement, or treat evidence from young children as if it automatically applies to teenagers.

Then do the part AI cannot responsibly do for you: open the papers, verify the claims, and discuss decisions with qualified professionals who understand the individual.

Do not paste private medical records, names, birth dates, school IDs, addresses, or other identifying information into a general-purpose AI tool.

AI Experiment: Audit the answer, not just the chatbot

Here is a simple experiment:

  1. Ask one autism-related question in the AI platform you normally use.

  2. Save the first answer without adding a follow-up prompt.

  3. Score it using the seven-part rubric above.

  4. Ask the AI to improve the answer based on the failed categories.

  5. Compare the first and second responses.

  6. Open every cited source and record whether it supports the claim.

The goal is not to crown one platform the permanent winner. The study found that platform patterns were largely stable across its two collection periods, but models, interfaces, and search features continue to change. The more durable skill is learning how to interrogate the answer in front of you.

Build With Me: Create an AI Health Answer Auditor in ChatGPT

The companion build for this article is a small, practical tool readers can create in ChatGPT. The AI Health Answer Auditor accepts an original question and an AI-generated answer, scores the response against the seven-part rubric, identifies red flags, checks available sources, and produces a safer revision.

It is not a diagnostic tool and it does not decide whether a treatment is appropriate. Its job is to slow the reader down long enough to evaluate the information before acting on it.

Copy-and-paste ChatGPT build prompt

Build an interactive tool called “AI Health Answer Auditor.” PURPOSE Help families, students, educators, and health-information consumers critically evaluate an AI-generated health or autism answer before relying on it. The tool is educational only. It must never diagnose, prescribe, recommend changing treatment, or replace a qualified professional. INPUTS 1. The user’s original question 2. The AI-generated answer to evaluate 3. Optional intended audience: parent/caregiver, autistic person, student, educator, clinician, or general reader 4. Optional age or population relevant to the question 5. Optional desired reading level PRIVACY GATE Before analysis, remind the user not to enter names, birth dates, addresses, school IDs, medical-record numbers, or other identifying information. If the pasted material appears to contain identifying details, pause and ask the user to remove them. AUDIT RUBRIC Score each dimension from 0 to 2: - Accurate: factual claims align with current credible evidence - Understandable: plain language, defined terms, manageable reading level - Actionable: concrete and reasonable next steps without diagnosing or prescribing - Evidence-supported: important claims have relevant, verifiable sources - Cautious: uncertainty, limitations, risks, and professional guidance are clear - Respectful: non-stigmatizing, neurodiversity-aware, autonomy-supporting language - Relevant: answer matches the age, population, setting, and question asked OUTPUT - A dashboard showing each score and a total out of 14 - A one-sentence explanation for every score - Reading-level estimate and difficult phrases to simplify - “What this answer did well” - “What is missing or potentially misleading” - Red flags: cure claims, certainty without evidence, blame, unsafe advice, discouraging evaluation, treatment changes, fabricated or irrelevant citations, or conclusions that exceed the population studied - A source table with claim, citation/link, source type, publication date, population studied, whether the source supports the claim, and verification status - A list of questions the user should ask a qualified professional - A revised answer written in clear, respectful language that preserves uncertainty SOURCE RULES When web access is available, open and verify each citation. Prefer original peer-reviewed studies, systematic reviews, clinical guidelines, and official public-health sources. Never invent a citation. If a source cannot be opened or verified, label it “unverified”—do not imply that it is reliable. Clearly separate research findings from inference or general educational suggestions. DESIGN Make the experience calm, accessible, and easy to scan. Use seven score cards, a 0–14 total gauge, green/amber/red indicators with text labels so color is not the only signal, and expandable evidence details. Include buttons for “Audit Answer,” “Improve Answer,” “Export Questions for My Appointment,” and “Start Over.” FINAL SAFETY NOTE Always end with: “This audit helps you evaluate information; it does not determine what is medically appropriate for a specific person. Verify important claims and discuss health decisions with a qualified professional.”

Final thought

AI can be a useful starting point for autism information. It can explain terminology, organize questions, summarize research, and help a family prepare for a conversation. But the quality of the interaction depends on more than the absence of a dangerous sentence.

The answer should be accurate enough to trust, clear enough to understand, practical enough to use, cautious enough to respect uncertainty, and humane enough to recognize the person behind the question.

That is not asking too much of health information. It is the minimum.

Call to action: Try the AI Health Answer Auditor with one answer you have already received. Then ask: Did the AI answer my question—or did it simply produce information near my question?

Sources

  1. Balikci, S. (2026). A comparative and temporal evaluation of autism information across ten AI platforms. Research in Autism, 136, 202965. DOI: 10.1016/j.reia.2026.202965

  2. National Library of Medicine. Evaluating Health Information.

  3. National Library of Medicine. Checklist: Evaluating Internet Health Information.

  4. Autistic Self Advocacy Network. What We Believe.

  5. Balikci, S. (2026). Quality of Autism Information Generated by Artificial Intelligence Tools: Implications for Paediatric Care. Journal of Paediatrics and Child Health.

Medical and Editorial Disclaimer

This article is for education and AI literacy. It does not provide medical advice, diagnosis, or treatment recommendations. AI output can be incomplete, outdated, or incorrect. Verify important health claims with original sources and consult an appropriately qualified professional for individual decisions.

Verification Flags for the Editor

  • Verified: Article title, author, journal, volume, article number, DOI, 10 named platforms, 15 questions, two collection periods, and evaluated dimensions.

  • Important qualifier: “No misinformation or unsafe advice” applies only to the response set evaluated in this study; it is not a universal safety finding about the platforms.

  • Update-sensitive: Platform features, free-access status, model versions, browsing capabilities, and default citation behavior can change after publication.

  • Terminology note: The source study uses “autism spectrum disorder/ASD.” This article also uses identity-respecting and neurodiversity-aware language while recognizing that individual language preferences vary.

  • Human review: Test all links and confirm the website’s final category/tag spelling before publishing.

Enjoyed this? Join Lab Notes.

One useful AI idea, one psychology insight, one build-in-progress, and one practical challenge — every week.