OpenAlex keywords used to be bad. Now they’re good. 419 million works in OpenAlex have new keywords, drawn from a vocabulary of 1.88 million. They’re 91% accurate, up from 44% for the old keywords.
The old keywords were guessed from each work’s topics, so they were often off. The paper announcing the first gravitational waves got “Black hole (networking)” and “Binary number.” The new tagger reads each work’s title and abstract and says what it’s about:
| Paper | Old keywords | New keywords |
|---|---|---|
| Observation of gravitational waves from a binary black hole merger (2016, here) | Black hole (networking), Binary number, Computer science | binary black hole mergers, LIGO, ringdown, tests of general relativity |
| Deep learning (2015, here) | Representation (politics), Layer (electronics) | backpropagation, convolutional neural networks, recurrent neural networks |
| Highly accurate protein structure prediction with AlphaFold (2021, here) | Computer science, Function (biology) | AlphaFold, protein structure prediction, protein folding problem, multi-sequence alignment |
How much better

Our judge is an AI model from another lab: OpenAI’s GPT-6 Astra, since our keywords were built with Anthropic’s Claude. It read 400 random works and rated each keyword on its own. You can see the results above. Shown both lists side by side, the judge preferred the new keywords for 395 of the 400 works.
The new keywords also match the keywords authors chose for their own papers twice as often as the old ones did, and they beat our previous best models on two standard keyphrase benchmarks. Every number is here.
The new keywords cover 89% of works, down from 95%. That’s on purpose: when a record doesn’t say what it’s about (a bare dataset entry, a table of contents), the tagger now returns nothing rather than a guess.
What they’re good for
- Finding what text search misses. A keyword finds the paper even when authors used different words for the same concepts. 133 million works with no abstract still get keywords.
- Mapping a field. Group any set of works by keyword to see what it’s made of, and by year to see what’s rising.
- Finding experts. Filter to a keyword and group by author; in our tests the top ten were real specialists on 24 of 24 queries.
- Searching across languages. Keywords are in English, but papers in any language get them: a Chinese paper on microplastics in soil carries the same
microplasticskeyword as an English one. About half of what the keywords add to a search on microplastics or remote work isn’t in English.

How to: find the papers your text search misses
Say you’re studying antimicrobial resistance. Just search: OpenAlex search now reads keywords along with titles and abstracts. When a phrase you type names a keyword, a paper matches if the phrase is in its title or abstract or the paper carries that keyword. A title and abstract search for the phrase finds 118,920 works; this one finds 241,167:
https://api.openalex.org/works?search.title_abstract_keywords="antimicrobial resistance"
The antimicrobial-resistance keyword brings in about 122,000 of them, and our judge rates 81% of those on topic: about 100,000 relevant works the phrase search missed, so you find about twice as many. Some say “antibiotic resistance” or “drug-resistant bacteria” instead. 47,269 have no abstract at all, and 74% of those are on topic, like “Increased resistance to antibiotics among microorganisms isolated in a general hospital” (1958, here) and “Lack of development of new antimicrobial drugs” (2005, here); no text search could find those.
On openalex.org there’s nothing to set: “Title, abstract, & keywords” is now the search box’s default, and the top results are reranked so the papers most clearly about your topic come first. Every other word you type still has to be in the text, so “remote work productivity” finds papers tagged remote work that mention productivity, never a paper tagged only productivity. In OQL it’s works where title-abstract-keywords has ("antimicrobial resistance").
Let your AI agent do a thorough search
The search catches the keywords your own words name. An AI agent can go further: find the keywords that mean your topic in other words, add them, and read the results to drop the strays. The /text/keywords endpoint takes any text, including a title or abstract, and returns the keywords that fit it, and an agent like Claude Code or Codex can use it for you. Tell it something like this:
Check help.openalex.org first, then use the OpenAlex API to find papers on <your topic>. Use the title, abstract and keywords search, and add OpenAlex keywords that describe my topic in other words, all in one query. Each paper should cover every part of my topic. Read the results as you go and keep only the papers that are really about my topic. Save the 200 most-cited of those as a CSV and show me the query and how many papers each part found.
My API key: <your key>
Keep the “check help.openalex.org” part: without it, agents work from what they remember about the OpenAlex API, which is out of date. We tested this prompt on fresh installs of Claude Code and Codex: every run found the title, abstract and keywords search on the help site and built one query with it, and on remote work and employee wellbeing, 99% of the papers in their final lists were at least partly on topic (about seven in ten clearly so). In the Claude app, set the chat to Auto (next to the model name, under the message box), or it will ask your permission before every search. You’ll need a free API key; the full recipe is here.
Or, if you use Claude, just ask. Add the free OpenAlex connector, now in Claude’s connector directory, and ask in one sentence: “Do a thorough search for open-access papers on how remote work affects employee wellbeing.” The connector searches titles, abstracts and keywords, adds keywords that mean your topic in other words, reranks the results so the most relevant come first, tells you how many works each part found, and hands back the query so you can rerun it.
The vocabulary keeps changing
This isn’t a fixed list. New fields show up every year, and we’ll add keywords for them as they do. We already merge keywords that mean the same thing (“Neanderthals” and “Neandertals”), and we’ll keep merging, splitting and adding as you tell us what’s wrong. Send us feedback on any keyword that misses.
Keywords are already linked to Wikidata: about 22% of keywords have a Wikidata id, and those cover about 56% of keyword uses. That makes it easy to join OpenAlex keywords to other data. Keyword search also knows 5 million synonyms, so searching keywords for “heart attack” finds myocardial infarction. Next, we’ll look at joining keywords to other outside vocabularies like MeSH, and at vector search, so that searching the keywords for “antibacterial resistance” finds antimicrobial resistance. We also want to assign each keyword to one or more of our topics, so keywords become part of the whole OpenAlex aboutness taxonomy.
Check our work
Everything is open: the code, prompts, benchmarks, vocabulary and model weights, with a release that lets you rerun the tagger yourself. The keywords help page has the details.
Edit, 3 October 2026: An earlier version of this post showed how to run a title and abstract search and a keyword search side by side and combine them by hand. A few hours after it went out, we shipped a search that does both at once (title, abstract and keywords, now the default on openalex.org), so we rewrote the how-to section and the AI prompt to use it.
Edit, 3 October 2026, later the same day: The OQL field in this post is now spelled title-abstract-keywords (it was title/abstract/keywords), because slashes already separate the parts of OpenAlex IDs. The old spelling still works.






