A 2D semantic map of the accepted NeurIPS 2026 papers, built from SPECTER2 embeddings of titles and abstracts. Nearby points are semantically similar papers, so a cluster is a topic.
Live site (after setup, see below): https://flecomet.github.io/neurips-explorer/
Derived from flecomet/cvpr-explorer, itself a fork of dataplayer12/cvpr-explorer. Original idea and design credit to @dataplayer12. Same license as upstream, see LICENSE.
The code is unit-tested; the paper data has not been generated yet.
public_submissions = false,
and every 2026 query returns 0 notes while 2025 returns papers). It also answers scripts running
on GitHub’s servers with a human-verification challenge. scrape.py (OpenReview) is kept for
when the papers are published./static/virtual/data/neurips-2026-orals-posters.json and ...-abstracts.json), which
scrape_site.py downloads. The abstract extraction from a paper page was checked against the real
site. The layout of the two JSON files was not known when the parser was written, so field names
are looked up from lists of likely candidates, and every run prints the structure it found.
python scrape_site.py --probe prints that structure without writing anything.?p=<paper id>, ?q=<search>, ?c=<topic>, ?color=<mode>./ search, Esc clear), usable on a phone.data.json, a few hundred KB gzipped) loads first, abstracts
(details.json) load afterwards.The pipeline is offline and the site is static (no backend, no API keys at serve time).
| Step | Script | Output |
|---|---|---|
| Fetch accepted papers from neurips.cc (or OpenReview) | scrape_site.py (scrape.py) |
data/neurips_2026_papers.json |
| Embed title + abstract with SPECTER2 | embed.py |
data/neurips_2026_specter2.npy (float16) |
| UMAP to 2D, HDBSCAN clusters, TF-IDF topic names, nearest neighbours | layout.py |
data/neurips_2026_layout.json |
| Merge into the site payload | build_site.py |
site/data.json, site/details.json |
site/index.html renders the payload client-side with plotly.js.
main.neurips.cc. It runs the whole pipeline
on a GitHub runner (embedding on CPU takes tens of minutes), commits data/, and starts the
Pages deployment. The Probe neurips.cc step prints what the site returned; if the scrape
step then fails, that output shows what the parser needs to change.committed uses data/neurips_2026_papers.json as it is in the repository.
openreview does not work from GitHub runners (human-verification challenge). Once OpenReview
publishes the papers, scrape from your own machine with python scrape.py, or from a browser:
open https://openreview.net, paste tools/browser_download.js into
the developer console, then run python scrape.py --raw neurips2026_raw.json. Commit the data
and run the workflow with source committed. The venue ids are in config.py.pip install -r requirements-pipeline.txt
python scrape_site.py # neurips.cc; for OpenReview: python scrape.py
python embed.py # GPU recommended, CPU works
python layout.py # --min-cluster-size 25 --n-neighbors 15
python build_site.py
python -m http.server -d site 8000 # http://localhost:8000
Tests: pip install -r requirements-dev.txt && python -m pytest.
layout.py is deterministic for fixed inputs (UMAP random_state=42) so reruns keep the map
stable. Changing the embeddings or paper set changes the map.--min-cluster-size 25. Lower it for
more, smaller topics.