Every work in OpenAlex now gets its topics from a new classifier. The topics themselves stay the same: same 4,516 topics, same names, same IDs, same 252 subfields, 26 fields and 4 domains. Your filters, dashboards and code keep working. What changes is which works are assigned to each topic, and the new assignments match what most readers expect: the topic fits what the title and abstract say.
It’s an incremental change. Institutions’ field profiles barely move, and university rankings by citation impact hardly change. Here’s what’s different, why, and what to do if you report on topics.
What’s different
Topics now match the title and abstract more often. Most people read a topic like a subject heading: they look at a paper’s title and abstract and expect the topic to describe them. The old topics often didn’t, and that confused people. One user wrote to us about a paper on universities: “it’s in ‘Diagnosis and Management of Kawasaki Disease’, though it’s about higher education in Australia.”
The mismatch came from how the old classifier worked. CWTS built the topics by clustering the citation network (our 2024 announcement), and the old classifier learned to predict which cluster a work belongs to, from its title, abstract, journal and references. For well-cited articles that works well, and a paper’s cluster usually matches its subject. Not always, and not always wrongly: a 2005 paper showing that hair follicle stem cells can repair severed nerves sat in “Hair Growth and Disorders”, because hair researchers cite it; the new classifier puts it in “Nerve injury and regeneration”. Both make sense. But for works with little to go on, the old classifier had to guess, and a few bugs made its guesses worse (below).
The new classifier reads each work’s title and abstract, in any language, and picks the topics whose subject matches. If the work has a venue, like a journal or a repository, it reads that name too, but it doesn’t need one, and it doesn’t use citations. So a dataset or a paper published yesterday gets topics as good as a well-cited article’s.
The gains are biggest where the old classifier had least to go on. We tested both classifiers on works that played no part in building the new one (how we tested):

It keeps what the clusters mean. We also ran the test the old classifier was built for: publications CWTS clustered, held out from the old classifier’s training, each scored against the cluster CWTS put it in. The new classifier never saw those labels, yet it matches CWTS as often as the old one does, 55% each, and has CWTS’s topic in its top three more often (73% against 70%) (details).
The bugs it fixes
From biggest to smallest:
- One topic held 20 million unrelated works. The old classifier gave most papers in Japanese, Chinese, Russian and Korean with no abstract the same three topics with the same scores, as if it couldn’t see their titles at all. The first was “Military Technology and Strategies“: 20.6 million works, 1 in 18 works with a topic. Now it holds 169,000 works, and they’re about the military.
- Works with no abstract got weak guesses. Four in ten works have no abstract. The old classifier leaned on abstracts and references; with only a title, its main topic matched the work 20% of the time. The new one: 75%.
- Datasets and other records that aren’t papers got no topic, or a stray one. Three in ten works are datasets, specimen records, proposals and the like. The old pipeline scored only papers, so most of these got no topic, and a classifier trained on journal articles misread the rest: 6% matched. The new one: 86%.
- Papers not in English got weaker topics. A quarter of works aren’t in English. The old classifier learned mostly from English papers (83% of its training set); 9% matched, partly because of the first bug. The new one: 72%.
- Records with nothing to classify got a topic anyway. A table of contents or a bare file name now gets no topic instead of a guess.
| Bug | Work | Old topic | New topic |
|---|---|---|---|
| One topic for 20 million works | Veselago’s 1967 paper on materials with negative ε and μ, in Russian, no abstract | Military Technology and Strategies | Metamaterials and Metasurfaces Applications |
| No abstract | A paper on hand-gesture recognition with millimeter-wave radar | Plant pathogens and resistance mechanisms | Hand Gesture Recognition Systems |
| Not a paper | A telescope proposal to map the tidal streams of globular clusters | Methane Hydrates | Galaxies: Formation, Evolution, Phenomena |
| Not in English | A Spanish-language history of Yiddish-speaking communists in Buenos Aires | Virology and Viral Diseases | Argentine historical studies |
Where the topics come from
CWTS at Leiden University built the topics in 2024 by clustering the citation network of OpenAlex works (our announcement). We built the classifier that assigns works to them. The classifier is what’s new; the clusters, their names and the hierarchy above them are untouched.
Your reports helped. Support tickets about topics showed us specific failure modes, like the 20-million-work bin, and the works in them became test cases. New AI models made the rest practical: a large model labeled two million works carefully, and we trained a small, fixed classifier on those labels. Its weights are open, so anyone can run it and get the same answer we do. The code, training labels and every test set are on GitHub.
If you report on topics
If you count your institution’s output by topic or field, or compare it with peers, here’s what changes and what to do.
Your field-level numbers barely move. We made before-and-after profiles for the 109 institutions that support OpenAlex. For the median one with at least 20,000 papers, its mix of fields is 97% the same and its mix of subfields 92% the same. Topics move more: 78% the same, and 7 of its top 10 topics stay in its top 10. Universities that publish a lot in languages other than English move most.
Rankings barely change. FWCI and citation percentiles compare a work with others in its subfield, so they shift when subfields shift. To see how much, we computed each university’s mean FWCI for its 2023–2025 papers both ways:

Across 3,621 universities, mean FWCI with the new topics tracks the old closely: the ranking barely changes (rank correlation 0.99), and the median university moves 4.5%. Universities that publish mostly in other languages stay on the line too, moving about as much as English-language universities with similar citation impact. Why anyone moves: the old catch-all topics sat in a few subfields, like Aerospace Engineering, and their millions of rarely cited works dragged those subfields’ baselines down. With fairer baselines, universities strong in those subfields slip a little, and universities strong in languages and the humanities rise a little.
Topics now cover every kind of work, not just papers. The old pipeline gave topics to papers; the new classifier gives them to datasets, specimen records and every other kind of work too. That’s part of our commitment to open science. For nearly 20 years, studies have shown that sharing data pays off (2007, 2013), and reformers of research assessment, from DORA to CoARA, ask that datasets and software count as research outputs. When datasets have topics, they show up next to the papers in their field. A few of these collections are huge: about 14 million fusion-experiment shot records now sit, correctly, in “Magnetic confinement fusion research” (39 million counting the expansion corpus, which API queries leave out by default), and millions of specimen records in botany and fungi topics. If your report should count papers only, filter by type (articles, reviews, books and chapters), and those topics look normal again.
Check your own numbers. The 109 profiles are on GitHub. For any other institution, group its works by field on the website or the API (MIT since 2015, by field) and compare with your last report.
See what changed. The change-set files list every topic, subfield and field with its works before and after, where each old topic’s works went, and, for each topic, how often an independent judge agreed with its works before and after.
Need the old topics? They’re kept. Every work’s old topics and scores, frozen on 5 October, are attached to the v2.0.0 release (12 files, 9 GB). The old classifier stays public too: its code is on GitHub and its weights and training data are on Zenodo, so you can keep running it. Our text-tagging endpoint also keeps it, at /text/topics?version=1, until 13 January 2027, so you have three months to compare the two side by side. Snapshot users: this change doesn’t touch updated_date, so reload fully from the 14 October snapshot.
Also new: topics and keywords are connected
Each keyword now shows the topics it’s most used in, with the share of its works in each (“Related topics”), and each topic shows its characteristic keywords. A keyword can belong to several topics, and we show how much. Explore topics, their keywords and the whole hierarchy in the aboutness viewer.
A month of aboutness improvements
Topics are the last of a run of changes to how OpenAlex describes what works are about. In each, your reports showed us what to fix. The middle column is our accuracy on the old system, then on the new one; the notes below the table say how we measured each.
| Change | Old → new | Links |
|---|---|---|
| Topics1 | 25% → 82% | code |
| SDGs2 | 29% → 70% | post, code |
| Keywords3 | 44% → 91% | post, code |
| Study designs4 | didn’t exist → 99.7% | post, code |
| Semantic search5 | 43% → 55% | list |
- Topics: share of 2,000 random works whose main topic matches the one two AI judges (Claude Opus 5.5 and GPT-6.1 Sol) each chose blind from all 4,516, counted where the judges agreed (1,503 works). A strict test built on reading the text; with no AI involved, the field matches the category authors chose for their arXiv papers 75% → 83%. Details.
- SDGs: F1, which balances precision (29% → 68%) and recall (31% → 72%), on 2,000 random works, against the goals two AI judges (Claude Opus 5.5 and GPT-6.1 Sol) chose, with Fable 5.1 breaking ties.
- Keywords: share of keywords an AI judge (GPT-6 Astra) rated accurate, on 400 random works.
- Study designs: precision of the RCT filter, the share of works it returns that are RCTs, judged by Claude Opus 5.5 on 8,308 papers and weighted to the whole index. OpenAlex had no study-design field before. The filter finds 69% of the RCTs PubMed tags in MEDLINE.
- Semantic search: how close the first page of results comes to the best possible order (nDCG), on 657 real searches graded blind by an AI model (Fable 5.1).
Thanks to everyone who sends us examples when something looks off. Keep them coming.






