This is a research search tool, not health information. What you see below are published abstracts and trial records returned by a similarity search — not reviewed, not ranked for quality, and not checked against current clinical consensus. Nothing here should inform a decision about a real child. For that, talk to a paediatrician or a developmental specialist.

Asking research questions about autism

A search over roughly 1,900 real research papers and clinical trial records, built to answer the questions families actually ask — and to show its working, so every answer traces back to the paper it came from.

Why this experiment

Autism is being identified at a rate that would have been unrecognisable twenty years ago. Part of that is broader diagnostic criteria and better recognition of children who were previously missed entirely; part of it may be a genuine increase. Researchers argue about the proportions, and the argument matters — but not to a family at the beginning of this.

What is striking, against that expansion, is how little has been settled underneath it. Decades of research have implicated hundreds of genes and a long list of prenatal and early-life factors, and produced no single cause. There is no treatment that addresses autism itself — what exists is support and intervention aimed at specific difficulties, which is a different thing and worth not confusing. An enormous amount is now known, and almost none of it resolves into the answer a parent is actually looking for.

So the questions families arrive with are rarely exotic. Is it genetic? Did something I do cause this? What does the evidence actually say about early intervention? What is being trialled right now, and where? Honest answers to all of them exist in published research — which is precisely where most people give up. It is written in unfamiliar vocabulary, scattered across tens of thousands of papers, and the readable summaries in between are too often produced by someone with something to sell.

This page is a small attempt at narrowing that gap: pointing a family at the actual literature on the question they asked, in the words the researchers used, with a link to every source. It answers nothing on its own authority, and it is not a substitute for a paediatrician or a developmental specialist.

Frequently asked questions

The questions people ask most often. Each runs the same search as the box below: a short summary of what the best-matching research says, with the sources it came from listed underneath.

Ask your own question

Anything you type is matched against the same 1,900 records. The citations in the summary link straight to the paper each claim came from.

Not connected yet. The search service has not been wired up, so free-text questions are unavailable. The questions above still work — they fall back to stored results, ranked but without a summary.

How this is builtthe sources, the models, the cost

Where the data comes from

Two public sources. No scraping, no paywalled content, nothing private, and no paid API.

SourceWhat it holdsHow it is accessed
PubMedUS National Library of Medicine ~1,500 abstracts of peer-reviewed autism research The free E-utilities HTTP API, in two calls: esearch returns the record IDs matching a search term, efetch returns those records as XML. No key, no account.
ClinicalTrials.govUS National Institutes of Health ~400 registered trials — what is being tried, by whom, and where Its public v2 JSON API, paged through with a cursor. No key, no account.

A scheduled GitHub Actions job re-fetches both sources monthly and rebuilds everything below, so the corpus does not quietly go stale.

Finding the right papers

Keyword search fails people here, because a worried parent and a journal abstract do not use the same words. "Will my son ever talk?" shares almost no vocabulary with a paper titled "Predictors of expressive language outcomes in minimally verbal children."

So each record is converted into a list of 384 numbers — an embedding — using BAAI/bge-small-en-v1.5, an open-weight model small enough to run on a laptop CPU. The numbers represent what the text means rather than which words it contains. A question is converted the same way, and the records whose numbers sit closest are returned, ranked by how close. That similarity score is printed on every result, because a number you can check is more honest than a confident answer you cannot.

There is no vector database. At 1,900 records the whole index is 2.8 MB — small enough to hold in memory and compare by brute force in well under a millisecond. A managed vector database would earn its keep at a scale this is nowhere near, and would add a paid dependency to something that currently costs nothing.

Why the second section needs a server

This site is static: GitHub Pages hands out files exactly as they sit in the repository and cannot run code. That is fine for the prepared questions, whose answers are computed offline and shipped as a data file. A question someone types cannot work that way, because nobody knew the question in advance — it has to be turned into numbers at the moment it is asked.

That runs on a Cloudflare Worker: a small piece of code that wakes on each request at whichever data centre is nearest the visitor. It embeds the question through @cf/baai/bge-small-en-v1.5 — the same model that built the index, which matters, because a question embedded by a different model lands in a different space and every score becomes meaningless while still looking plausible. It ranks the 1,900 records, then passes the best four to @cf/mistral/mistral-7b-instruct-v0.2-lora to write the summary.

Both models run on Cloudflare's free tier, which allows 10,000 neurons a day — a question costs a small fraction of one percent of that. The whole thing, end to end, runs at no cost.

What the model is and is not allowed to do

The summariser sees only the retrieved records, never its own training knowledge. It is instructed to put a citation marker after every factual sentence, to say plainly when the sources do not answer the question, and never to give medical advice or suggest a diagnosis. If the search finds nothing above a relevance threshold, the model is not called at all — asking it to answer from no sources is inviting it to invent one.

Those are instructions, and instructions get ignored. So the page checks too: a citation is only turned into a link when it points at a source that is genuinely on the page, and the summary is styled distinctly from the result cards below it — those are quoted verbatim from real papers, the summary was written by a model, and a reader should never have to guess which is which.

What is still missing

The corpus is deliberately narrow: research abstracts and trials, nothing else. The obvious additions are prevalence and demographic figures from the CDC's monitoring network, each carrying its own citation, and the trial locations already sitting unused in the data, which could become a map of who is doing this work and where.