Find research gaps in the corpus you already collected
Upload the spreadsheet from your search, or the article PDFs, and see which keywords are rare, which pairs never appear together, and which terms grew or faded.
Free credits · No credit card · Instant access What changed recently
What this solves
The Research Gap Detector reads the corpus you upload: a spreadsheet from your bibliographic search, PDFs of articles, or both. It counts author keywords, lists the ones that appear only a few times, finds frequent pairs that never occur in the same record, and compares the newest year and the three before it against the earlier ones. What is a real gap is your call.
How it works
-
Upload your corpus
Send the .xlsx exported from your search, the PDFs of the articles, or both. Fill in the research area if you want the analysis contextualized.
-
Run the analysis
The tool counts author keywords, crosses the frequent pairs and compares periods. Keep the AI option on to also get an executive summary and research problems.
-
Read the result and decide
Read the executive summary, the trend reading and the suggested research problems on screen. Export to JSON or Excel to go through the full lists of underexplored keywords and unexplored pairs (the CSV is a shorter summary), and take what holds up in your field.
Method
What happens to your file, step by step, and where the result stops being a decision made by a model.
How it works inside
For Excel, the first sheet is read (if it has fewer than two columns, the next sheet with more columns is used) and keywords come only from columns named 'Author Keywords', 'Keywords' or 'DE', split on ';' and ',', lowercased and counted. Underexplored themes are keywords with frequency 1 to 5. Unexplored pairs are every combination of the 30 most frequent keywords minus pairs that already co-occur, ranked by mean frequency. Trends split records at (max year minus 3): emerging above +50 percent growth, declining below -30 percent. PDFs follow a separate, model-driven path.
Where AI is used, and where it is not
In the Excel path the counting, the 1 to 5 frequency window, the pair enumeration and the growth thresholds are computed by fixed code, with no AI involved. GPT-5-mini receives those numbers and writes only the executive summary, the trend narrative and five research questions. In the PDF path the model decides the substance: it extracts title, year and keywords from each file and names the gaps itself. If the call fails, canned fallback text is stored.
What it accepts
- One .xlsx or .xls file plus any number of .pdf files, sent in the same request (whole request capped at 150 MB by the app). If you send two spreadsheets, only the first one is read.
- The spreadsheet only yields keywords from a column named exactly 'Author Keywords', 'Keywords' or 'DE'. Without one of them, the analysis runs on zero keywords.
- The year column is found by name ('Year', 'PY', 'Publication Year', 'Ano' and variants) or, failing that, by scanning for a column where at least 50 percent of a 50-row sample are integers between 1900 and 2100.
- Optional: research area (240 characters in the form), a custom title, and an AI toggle that starts on.
What you get back
- The result screen shows the totals, the executive summary, the AI trend reading and the research problems. The full lists (top 30 keywords, up to 50 underexplored themes, up to 50 unexplored pairs, and emerging, declining and stable lists capped at 30 items each) are stored with the analysis. The JSON download carries all of them; the Excel carries the lists but leaves the stable trends out, and the CSV is a summary cut to 30 themes and 20 pairs.
- Downloads: JSON with the complete result object, XLSX with up to six sheets (summary, frequent keywords, underexplored themes, unexplored combinations, trends, AI suggestions) and a semicolon-delimited CSV.
- AI block: executive summary, 6 to 8 trend items, and 5 research problems, each with a suggested methodology and variables. Regenerating the problems is a separate paid action (10 credits against the 20 of the analysis).
- The full result is saved to your history at its own address and can be reopened, renamed or removed from the history.
What this tool does not do
- No stemming, lemmatization or synonym merging. 'machine learning' and 'machine-learning' count as two keywords, and a file that has both an 'Author Keywords' and a 'Keywords' column counts the same term twice.
- 'Underexplored' means literally 'appears 1 to 5 times in the file you uploaded'. The tool queries no external database and has no idea whether the theme is well explored elsewhere.
- An unexplored pair only means the two keywords were not seen together in the same record, among the 30 most frequent keywords, checking the first 10 alphabetically sorted keywords per record. It is not evidence that nobody has studied the two topics together.
- Without a detectable year column, trend analysis returns empty lists and the model is explicitly told the trends have no years behind them.
- It does not judge whether a gap matters, is fundable or is publishable, and it does not check whether the suggested research questions have already been answered.
- Plan
- Pro
- Cost
- 20 credits
What you get
-
Underexplored keywords
Author keywords that appear only a few times in your corpus, listed with their counts, so you can see how thin the coverage is.
-
Pairs never studied together
It takes the most frequent keywords and lists the pairs that never show up in the same record, ranked by how common each term is alone.
-
Emerging and declining terms
With a year column in your file, it compares the newest year and the three before it against the earlier ones and separates what grew from what faded.
-
Everything exportable
Download the analysis as an Excel workbook with one sheet per section, or as CSV or JSON. Past analyses stay in the sidebar history.
Changelog
Every line below is a change that actually shipped, dated by the day it went out.
-
Latest
- Regenerating research problems announced 5 credits for an operation that costs 10. The refusal sent when credits are missing now carries the real cost from the price table.
Show 5 earlier updates
-
- When the AI provider fails, you read a short explanation instead of the raw technical error.
- Authentication trouble and quota trouble now carry different texts, because they ask for different actions.
- The fallback path that still returned a raw error was covered as well.
-
- With a project open, title and research area come prefilled from that project.
- Analyses saved with a project open are now collected inside that project.
-
- The tool page now requires login and a plan that includes the gap detector.
- Interface texts and messages go through translation, so English readers see English.
-
- Regenerating research problems went from 2 to 5 credits, with the on-screen text updated.
- Running an analysis now checks that you are logged in before starting.
-
- The analysis and the generated research problems now follow the language of the interface.
Questions and answers
Does it prove that a gap exists?
No. It counts what is in the files you sent. A rare keyword can be an open niche or a path the field already abandoned. Reading the papers and deciding stays with you.
What does the file need to have?
A spreadsheet in .xlsx or .xls with an author keywords column (Author Keywords, Keywords or DE), and a year column for the trend analysis. PDFs of articles work too, alone or alongside. Each analysis costs 20 credits on the Pro plan.
Does the AI invent the research problems?
It writes the summary and the suggested questions from the counts computed in your own file, and it is instructed not to invent growth percentages. Even so, check every suggestion against your reading before adopting it.
Is my data used to train models?
The extracted text goes to OpenAI on a paid API tier, whose published policy states that API content is not used for training. Your analyses stay in your history until you delete them.
How to cite this tool
Used it in your research? Here is the reference, already filled in with the version you are looking at and today's access date.
ABNT (NBR 6023)
LESSA, P. W. B. Research Gap Detector. Versão 2026.08. [S. l.]: Xplore Dados, 2026. Disponível em: https://xploredados.com/en/tool/research-gap-detector. Acesso em: 21 ago. 2026.
APA 7
Lessa, P. W. B. (2026). Research Gap Detector (Version 2026.08) [Computer software]. Xplore Dados. https://xploredados.com/en/tool/research-gap-detector
BibTeX
@software{xploredados_research_gap_detector_2026,
author = {Lessa, Patrick Wendell Barbosa},
title = {Research Gap Detector},
organization = {Xplore Dados},
version = {2026.08},
year = {2026},
url = {https://xploredados.com/en/tool/research-gap-detector},
urldate = {2026-08-21}
}
Start with the free credits
Create an account and test the tools before deciding on a plan.
Create free account