Talos brings continuous genomic reanalysis to nearly 5,000 unsolved cases

Talos brings continuous genomic reanalysis to nearly 5,000 unsolved cases

Concept 1: Why Genomic Diagnosis Is Hard — The "Incomplete Knowledge" Problem

What is genomic testing?

Genomic testing reads a patient's DNA to look for variants — differences in their genetic code that might cause disease.

The core problem

Even after testing, more than half of rare disease patients remain undiagnosed. This is not because the test failed to read the DNA correctly. The DNA data is there. The problem is that:

We don't yet know what all of it means.

Every year, researchers discover:

  • New gene–disease associations (e.g., "we just learned that mutations in Gene X cause Disease Y")
  • New variant classifications (e.g., "this specific DNA change, once considered harmless, is now known to be disease-causing")

The key insight

Unlike a blood test or an X-ray, genomic data doesn't expire. It can be stored and re-examined later. This means:

A diagnosis that was impossible to make in 2020 might be straightforward to make in 2025 — using the exact same DNA data.


Concept 2: Reanalysis — The Logical Solution

What is genomic reanalysis?

Reanalysis means re-running the interpretation of a patient's already-sequenced DNA against newer, updated scientific knowledge.

Think of it like this:

  • The DNA sequence = a book written in a language
  • Scientific knowledge = the dictionary for that language
  • As the dictionary grows, you can understand more of the book

Does it work?

Yes. A meta-analysis of ~9,500 undiagnosed patients found that reanalysis increased diagnostic yield by ~10% over roughly two years.

So why isn't everyone doing it?

Because today, reanalysis is overwhelmingly manual. It requires:

  • Motivated clinicians
  • Scarce laboratory staff
  • Consistent funding/reimbursement

The result: the vast majority of stored genomes are never revisited, even as the data keeps accumulating.


Concept 3: The Core Trade-Off in Automated Reanalysis

Why is automation hard?

Any automated system must balance four competing pressures:

TermWhat it meansWhy it matters
SensitivityAbility to catch true diagnosesMissing a diagnosis = patient stays undiagnosed
SpecificityAbility to avoid false alarmsToo many false positives = analysts overwhelmed
Candidate variants per patientHow many results a human must reviewDirectly determines whether the system is sustainable
Frequency of reanalysisHow often you re-run the analysisMore frequent = faster diagnoses, but more work

The real bottleneck

The limiting factor in real-world genomic reanalysis is not the algorithm's ability to find variants. It is human expert review time. An automated tool that flags 50 candidates per patient is not useful if analysts can only realistically review 1–5.

This is the central design challenge Talos was built to solve.


Concept 4: How Talos Works — The Pipeline

What is Talos?

Talos is an open-source, automated tool that re-interprets a patient's existing variant data against the latest scientific knowledge each time it runs.

Step-by-step pipeline

Stage 1 — Static Annotation (collect unchanging information)

  • Gather the patient's existing variant calls (from prior sequencing)
  • Record fixed biological facts about each variant (e.g., which gene it's in, what type of change it is)

Stage 2 — Dynamic Annotation + Prioritization (apply up-to-date knowledge) Talos draws on two continuously updated public databases:

  • PanelApp Australia → gene–disease relationships and inheritance patterns
  • ClinVar → variant-level pathogenicity classifications

It then applies a variant-prioritization algorithm designed to surface variants most likely to meet clinical reporting standards (ACMG/AMP criteria).

It also uses:

  • Family structure (e.g., did the variant arise de novo — new in the child, not inherited?)
  • Mode of inheritance (e.g., does the disease require one or two faulty copies?)
  • Patient phenotype (e.g., does the patient have heart problems? Filter for cardiac genes)

Stage 3 — Reporting

  • Surface a small, high-confidence set of variants to clinicians
  • Include supporting evidence so reviewers can act quickly

What types of variants can Talos handle?

  • Single-nucleotide variants (SNVs) — single letter changes in DNA
  • Small insertions/deletions (indels)
  • Copy number variants (CNVs) — extra or missing chunks of DNA
  • Large structural variants — major rearrangements of chromosomes

Concept 5: Two Key Design Choices That Make Talos Different

Choice 1: Be conservative (optimize for specificity, not sensitivity)

Most tools return a long ranked list of candidates. Talos deliberately returns a short, high-confidence set.

  • Median output: 1.3 candidate variants per patient
  • This respects the real bottleneck: analyst time

Think of it like a search engine that shows you 3 highly relevant results vs. one that shows 3,000 results ranked by relevance. If you only have time to read 3, the first approach wins.

Choice 2: On repeat runs, only show new findings

On each iterative cycle, Talos only flags variants whose supporting evidence has changed since the last run.

This means analysts aren't re-reviewing the same variants over and over. They only see what is genuinely new.

Result: In monthly cycles, analysts needed to review only 1 new variant per 200 patients — making frequent reanalysis sustainable.


Concept 6: Validation — Does It Actually Work?

Benchmark cohort 1: Australian Acute Care Genomics (ACG)

  • Critically ill infants and children
  • Talos recovered 90% of in-scope diagnoses
  • Median of 1.3 candidate variants per family

Benchmark cohort 2: Rare Genomes Project (RGP) — USA

  • Families with prior uninformative clinical testing
  • Patients ranging up to 82 years old
  • Talos recovered 87% of in-scope diagnoses (47 of 54)
  • Same operating point: 1.3 candidate variants per trio

The fact that performance held across two very different cohorts demonstrates generalizability — the tool isn't just tuned to one specific population.

Head-to-head vs. Exomiser (a widely used tool)

  • Overall sensitivity: statistically similar
  • But at a realistic review budget (top 1–5 variants), Talos was significantly better
  • The two tools flagged different variants → they are complementary, not competing

Concept 7: Real-World Deployment — The 5,000-Patient Cohort

The experiment

Talos was deployed on 4,735 previously undiagnosed patients from Australian Genomics research studies and a diagnostic laboratory.

Results

  • 241 new diagnoses in 238 individuals
  • 5.1% additional diagnostic yield
  • Every flagged variant was subsequently confirmed as pathogenic or likely pathogenic by accredited labs

Where did the diagnoses come from?

Source% of diagnosesWhat this means
New gene–disease relationships32%Science discovered new gene links after original test
New variant-level evidence22%A variant was reclassified as disease-causing
Improved filtering/analysis45%Better methods, broader variant types (CNVs), refined phenotype filters

Notable finding

59% of new gene–disease diagnoses were not yet in OMIM (the standard reference database) at the time of reanalysis. This highlights the value of using a rapidly updated resource like PanelApp Australia rather than slower-updating databases.


Concept 8: Continuous Reanalysis — From One-Off Event to Ongoing Program

The iterative experiment

Talos was run for 29 monthly cycles on the same cohort.

Key findings

Finding 1: Most value comes on the first pass, but iteration still matters

  • 92% of diagnoses came on the first run
  • But later cycles captured diagnoses that only became possible as new science emerged

Finding 2: Iteration is sustainable

  • Later cycles flagged only 1 new variant per 200 patients
  • This is a workload that clinical teams can realistically absorb

Finding 3: Speed from discovery to diagnosis

  • Average of 32 days between new knowledge appearing in a public database and a patient receiving a diagnosis
  • Fastest case: 1 day

Compare this to the traditional model, where patients might wait years — or forever — for someone to manually revisit their file.

Cost

  • Annotating 1,000 genomes: ~$11
  • Monthly reanalysis pass: a few cents per cohort

This makes continuous reanalysis economically viable at scale.


Concept 9: The Bigger Picture — What Talos Represents

A paradigm shift

Talos reframes genomic reanalysis from:

❌ A rare, labor-intensive, inconsistently funded event

To:

✅ A continuous, automated program that keeps pace with science

The key enablers

  1. Openly shared, frequently updated databases (PanelApp Australia, ClinVar) turn global scientific progress into individual diagnoses
  2. Conservative design respects the human bottleneck
  3. Incremental reporting makes iteration sustainable
  4. Open-source + cloud deployment makes it accessible to health systems worldwide

Looking ahead

The authors anticipate that AI models for predicting the consequences of genetic variation will further enhance reanalysis — as these models improve, they can be plugged into the Talos pipeline to catch even more diagnoses.


Summary: The Conceptual Arc

Incomplete genomic knowledge
        ↓
Patients remain undiagnosed after first test
        ↓
Reanalysis is the solution — but manual reanalysis doesn't scale
        ↓
Automation is needed — but must balance sensitivity, specificity, and reviewer burden
        ↓
Talos solves this by being conservative + incremental
        ↓
Validated at scale: 241 new diagnoses, 5.1% yield, 32-day average turnaround
        ↓
Continuous reanalysis becomes a sustainable, affordable clinical program

The central lesson of this article is that the value of genomic data doesn't end at the first test — and with the right automated tools, health systems can continuously convert accumulating scientific knowledge into diagnoses for patients who have been waiting, sometimes for decades.

More to study