LacuņaEnter system
Research gap and discovery engine

Every field has a map.
We trace what is left uncharted.

Built for researchers who want to work where science needs them most, indexing 0 publications to expose critical, overlooked gaps.

Enter systemWhy it exists
Top 25 of 250+ at NeuroLogic ’26 Global NLP Datathon
0
papers indexed
0
languages tagged
0%
study English or peers only
79 languages, by coverage
Centre is saturated, rim is neglected. Drag any node.
SaturatedEmergingNeglected
01 · The problem

Finding a gap means reading for months, and still guessing

A researcher choosing a direction has to hold an entire literature in their head, then bet that the absence they noticed is real and not simply a subject nobody needs.

0%
of papers study English or its peers only
Out of every paper in the corpus carrying a language tag, roughly half never leave the highest-resource tier.
0M
speakers per paper for Saraiki
Tier 0, and the whole indexed literature on it amounts to 2 papers.
0/80
pairings with no paper at all
In a small sample of ten languages against eight common tasks, this many combinations have never been published on.
Three failures compound
  1. It does not scale. A literature review is bounded by how much one person can read, so the map is always a decade and a subfield wide at best.
  2. An absence is ambiguous. Nothing published on a topic can mean it is untouched, or that it is a non-problem. Reading alone cannot separate the two.
  3. The bias is invisible from inside. If your field studies English, the shape of what it skips never appears in the papers you read, only in the ones that were never written.
Speakers carried per indexed paper

Millions of people who speak the language, divided by how much research exists on it.

SaraikiT0
13M · 2 papers
JavaneseT1
7M · 11 papers
MaithiliT0
6M · 6 papers
FulfuldeT0
5M · 5 papers
SundaneseT1
5M · 7 papers
LaoT1
3M · 10 papers
BhojpuriT1
3M · 19 papers

English holds 3,577 papers in the same corpus.

Afieldrecordswhatitstudied.Itneverrecordswhatitskipped,sotheabsencehastobemeasuredinsteadofread.

02 · The vision

Treat the blank space as a measurable object

If a field's coverage can be counted, then so can its absences. That turns choosing a research direction from an act of intuition into an argument with evidence attached.

01
Absence is data

An empty cell is not a missing record. It is a finding, and it should be scored, ranked and defended like any other result.

02
A gap must be argued

Nothing published is only interesting when comparable languages have solved the same task. That adjacency is what separates an opportunity from a non-problem.

03
Every number is checkable

A statistic with no path back to its records is a claim, not evidence. Each figure opens the papers that produced it.

04
No black box

Tags come from an explicit gazetteer, not model inference, so the same query returns the same answer and a reader can audit the reasoning.

03 · What it solves

The map, with its holes drawn in

Ten languages against eight common tasks. 6 of these 80 pairings have no indexed paper, and the hatched squares are exactly where a new contribution has no competition.

Translation
NER
Sentiment
Toxicity
QA
Speech
Summarisation
Dialogue
EnglishT5
Mandarin ChineseT5
ArabicT5
HindiT4
BengaliT3
UrduT2
SwahiliT2
AmharicT2
YorubaT2
SindhiT1
no papera fewwell covered
04 · How it works

Four moves, from a phrase to a defensible shortlist

Each step is visible in the interface, and each one can be interrogated rather than taken on trust.

05 · How it is built

A pipeline you can audit end to end

Deterministic from source to score. No model sits in the chain, so the same question always returns the same answer.

01
Sources
ACL Anthology bulk export
OpenAlex sweep
130,930
entries scanned
02
Relevance gate
Low-resource or multilingual
Abstract required
16,612
papers kept
03
Gazetteer enrichment
Languages, tasks, methods
Datasets, code-mixing
79
languages tagged
04
Retrieval
BM25 over title and abstract
Taxonomy expansion
57,730
distinct terms
05
Gap engine
Coverage matrix, momentum
Five-component scoring
0–100
opportunity score
No model sits anywhere in this chain. Every tag is a literal string match against an explicit gazetteer, which is what makes each number reproducible from the source text and traceable back to the papers behind it.
06 · Start

Pick a starting point, or bring your own question.

Every run produces a ranked set of gaps, generated research questions, a downloadable brief and a link that reopens the exact analysis.

Enter system