
Most clinical diagnoses can be made through language-based interviews β a doctor simply asking you questions about how you feel. However, these interviews have real-world barriers:
Current AI language models (LMs) have been tested on curated, highly detailed medical case studies β essentially "textbook perfect" patient descriptions. But real patients:
Key Insight: There is a significant gap between how AI performs on clean test cases versus how it would perform with real everyday patients.
A differential diagnosis (DDx) is a ranked list of possible diagnoses that could explain a patient's symptoms.
If you report fever, sore throat, and fatigue, a DDx might include:
Because symptoms overlap across many conditions. A clinician β or AI β narrows down possibilities through follow-up questions before arriving at a most likely diagnosis.
Key Metric Used: Top-5 Accuracy β whether the true diagnosis (later confirmed by a real doctor) appears anywhere in the AI's list of 5 candidates.
SymptomAI is a conversational AI agent built on Gemini Flash 2.0 that:
A panel of three board-certified clinicians:
β οΈ Important Note: All AI-generated diagnoses were for research purposes only β not official medical assessments.
Participants were randomly assigned to one of five agent types, each with a different interview style:
| Study Arm | Description |
|---|---|
| Dynamic Live | Fully unrestricted follow-up questions in real time |
| Dynamic Final | Unrestricted questions, summarized at the end |
| Fixed Canonical | Standard medical school history-taking questions |
| Flexible Canonical | Standard questions with some flexibility |
| Base (Control) | No agent guidance β user drives the conversation |
All four agent-driven approaches significantly outperformed the Base condition.
Lesson: When AI actively asks follow-up questions, diagnostic accuracy improves dramatically compared to just letting users type whatever they want.
In more than 50% of cases, clinicians ranked SymptomAI's DDx as better than those provided by other clinicians.
SymptomAI's DDx more frequently contained the true diagnosis (as confirmed by the participant's personal doctor) compared to clinician-generated DDx.
SymptomAI performed best relative to clinicians in cases where clinicians themselves felt least confident β suggesting AI may add the most value precisely where human uncertainty is highest.
Big Picture: SymptomAI performed at or above the level of board-certified clinicians in this research setting.
If SymptomAI correctly identifies someone as having a respiratory infection, we should be able to see physical evidence of that illness in their body data β even before they reported symptoms.
Participants diagnosed by SymptomAI with respiratory infections showed:
Biosignal Change
β
| π΄ Infected (respiratory)
| β±
| β±
|____β±___________________________
| β¬ Baseline (all others)
|
ββββββββββββββββββββββββββ
Day -30 Day -15 Day 0 (SymptomAI conversation)
Why This Matters: The wearable data independently confirms SymptomAI's diagnoses β providing objective, physiological evidence that the AI got it right.
To study health patterns across large populations, researchers need clinical-quality labels (i.e., someone needs to diagnose each person). This is:
By accurately diagnosing thousands of people automatically, SymptomAI can:
Analogy: It's like having a tireless clinician who can interview and diagnose millions of people, then hand that data to researchers to find patterns no single hospital could ever study.
The researchers openly acknowledge several important caveats:
PROBLEM β Symptom assessment has barriers; AI tested only on clean data
SOLUTION β SymptomAI: conversational AI that interviews real patients
METHOD β 13,917 participants, 5 agent types, clinician evaluation panel
RESULT 1 β SymptomAI DDx preferred over clinicians' in 50%+ of cases
RESULT 2 β Wearable biosignals independently validate AI diagnoses
RESULT 3 β Active follow-up questioning is key to accuracy
FUTURE β Population-scale health research now becomes feasible
Core Takeaway: SymptomAI demonstrates that conversational AI can conduct real-world symptom interviews at or above clinician-level accuracy β and when paired with wearable data, opens entirely new possibilities for large-scale health research.