How “RAG over the whole collection” works
This is a two-stage retrieval hierarchy — it never embeds the whole hub (petabytes, impossible client-side), it instead retrieves the right datasets first, then retrieves inside them:
| Stage | What happens |
| 1 · Term extraction | The question is broken into weighted keywords (stopwords removed). |
| 2 · Hub discovery | Those terms hit the HF Hub search API — candidates ranked by downloads/likes. |
| 3 · Relevance hierarchy | Each candidate dataset is tiered: tier 1 id/tag match → tier 2 description match → tier 3 broad/popular — plus a meta-score. |
| 4 · Content probe | The top candidates’ real rows are fetched and lexically scored against the question (TF-style). |
| 5 · Prune & re-rank | Datasets with no relevant rows are dropped; chunks from all survivors are merged and ranked. |
| 6 · Synthesize | Extractive answer by default (no key needed) or model-generated (&model=…) with row citations. |
⚡ Everything runs in your browser — the Space is pure static (HTML/JS). The only network calls go to HF’s public, free APIs. No server, no billing, no API key required.