About LibCorpus

What the collection is, how a paper gets in, where each panel's data comes from, and how to use it.

LibCorpus hosts the numerical data behind figures from open-access research papers, re-rendered as interactive plots and served as machine-readable files. The unit of the collection is the figure panel: each one carries its own data and its license, and can be downloaded exactly as served.

LibCorpus is built and maintained by Dr. Jabir Chathanathil (LinkedIn ORCID), a physicist working in quantum sensing and control.

For what the collection currently holds, panel by panel, see the corpus statistics.

What gets included

Papers arrive three ways: through automated discovery of eligible open-access papers, by direct selection, and through submissions from authors. In every case the decision to include a paper is made by a person, and every paper is reviewed before publication. A paper qualifies when:

The two conditions above decide what qualifies, not the field. The collection began with quantum science and technology, and today that is what it holds: quantum sensing, precision measurement and adjacent atomic and optical physics. A paper from any field that meets both conditions can be included, and authors of any open-access paper can submit their own data directly.

Where the data comes from

Each panel identifies how its data was obtained, and the same statement travels in the header of every download. We distinguish six provenance routes, and the corpus statistics page counts how many panels come by each:

A single figure may mix routes; the panel's note says which.

What a hosted figure is, and is not

Hosted panels are re-rendered from the hosted data for interactive viewing, keeping the published figure's axes, ranges and content, but they are not the original graphic. The paper PDF, embedded beside every figure, remains the reference. When a journal has accepted a paper but not yet published the version of record, and the accepted manuscript is openly licensed, the panels are built from the manuscript and the version of record replaces it once it appears.

Not everything is hosted. Illustrative figures with no extractable data (setup schematics, level diagrams) are listed but marked "in PDF". Figures whose data could not be obtained, or could not be matched to the figure unambiguously, say so in exactly those words. We do not silently omit figures.

Curation and review

Extraction is AI-assisted: an automated pipeline fetches the deposit or runs the authors' code and builds the data files and metadata for each panel. Before anything becomes public, each panel passes two checks:

  1. Human review. In a review desk, every rendered panel is compared with the printed figure and approved, sent back for changes, or dropped.
  2. Validation. Automated checks verify that every published panel's data and metadata are present, parse cleanly, and that the paper's stated data-point count matches its package.

Only approved panels reach the site. A paper is released when its review is complete: every hosted panel approved, and any panel left out recorded on the paper's figure list.

Review is a good-faith check, not a guarantee. The published paper and the authors' own deposit remain the authoritative record; the hosted data is provided as-is, and where the two ever disagree, the original sources govern. When we learn of an error, we correct the panel or take it down.

Corrections

To report an error in any figure or dataset, email contact@libcorpus.org with the paper ID and panel. Confirmed errors are corrected, or the panel is taken down; a paper can also be delisted entirely, which removes it from the site.

Every package carries a version number, bumped whenever its data or metadata are re-released, and every download header records it, so a corrected panel is identifiable from the file alone. Authors are notified when their paper goes live, and an author can ask at any time for a panel or a paper to be amended or taken down.

Licensing

Each paper keeps its own license, shown as a badge on its page. It is the license of the version hosted here, which for some papers is the arXiv preprint rather than the journal article. When the hosted data carries different terms than the article, both are shown: a CC BY paper with a CC BY-SA deposit, for example, shows both licenses, and the stricter terms govern the downloads. Data that authors shared with us directly, without a license of their own, is hosted with their permission and no license is granted for it; its page says so, and reuse beyond citation needs the authors' agreement. Every downloaded CSV carries its provenance and its license or terms in its header.

Using and citing the data

Every download may be reused under the license or terms shown on its paper's page, which are also written into the file's header together with the paper's citation and DOI. When you use hosted data, cite the original paper, and give the panel's URL so a reader can find the exact values you used. Ready-made citations, as BibTeX or CSL-JSON, are available for every paper and every panel through the API (see For machines); their note names the LibCorpus page and the package version.

For machines

Each panel's download is a CSV of the plotted values, with a second CSV when the panel overlays a fit or a model. Every CSV begins with a header naming the panel and its page, the paper, the version hosted and its license, the data's license or terms, the package version and the provenance route. What each column holds, how the file is laid out and which columns form each curve are in the panel's record in the API. Panel pages have stable URLs.

Every panel's files are also served as plain files at stable addresses of one shape, /api/v1/papers/<paper>/figures/<figure>/panels/<panel>/files/<name>: data.csv for the plotted values, and beside it any overlay and the plotting metadata. CSV files carry the header above; add ?header=0 for the stored file without it. They answer any origin, so a notebook, a script or an AI assistant can read a panel with one request, for example pandas.read_csv(url, comment="#"). The catalogue and every paper, panel, About and statistics page carry their content in plain HTML as well, so a crawler or an assistant that does not run JavaScript reads the same text and links a reader sees. These addresses do not change.

For language models there is a text view of everything: /llms.txt describes the library and lists its papers, and any paper, panel, About or statistics page is available as Markdown by adding .md to its address (or asking for text/markdown): the page as text, with each panel's description added. A DOI or arXiv id resolves to its paper at /doi/<doi> and /arxiv/<id>.

The JSON API is documented at /api/docs (OpenAPI document at /openapi.json): one call for a paper with every figure and panel, one for a panel with its columns, files and hashes, a search over papers and panel descriptions, and citations as BibTeX or CSL-JSON. The v1 addresses and fields do not change; fields may be added.

An AI assistant can use LibCorpus as a tool directly: add https://libcorpus.org/mcp as a remote MCP server (Model Context Protocol) in Claude, ChatGPT, Cursor or any other MCP client, with no key. It then has tools to search the corpus, read a paper or a panel, pull a panel's data and cite the paper. Step-by-step instructions for each client are on the API reference.

Contact

Questions, corrections, and submissions: contact@libcorpus.org, or use the submission form.