A short, honest account of where the index comes from and how “open access” is decided — written for anyone deciding whether to trust a link before clicking it.
Cortexa pulls from three places, rebuilt daily: arXiv (via its public Atom API, across robotics, ML, AI, computer vision, and related categories), CrossRef (covering MDPI, individually open-access IEEE Transactions articles, and other open-access publishers), and OpenAlex (a broad scholarly index, used both for general coverage and for excluding results already covered by arXiv so the sources stay complementary rather than redundant).
“Open access” isn't one uniform thing across publishers. MDPI is 100% open access by policy, so every MDPI paper qualifies automatically. Everywhere else — IEEE, Springer, Elsevier, and other publishers — a paper only gets indexed if its own record carries a verified Creative Commons license, regardless of which journal it technically appeared in. That's what lets Cortexa include individually open-access articles from otherwise-paywalled IEEE Transactions journals, without misrepresenting the paywalled majority of that same journal as free.
A paper never gets marked open access just because it showed up in an open-access- sounding search result — the license check happens per-article, every time.
Publisher APIs sometimes list a metadata endpoint as the “PDF link” instead of the actual file — clicking it opens raw XML instead of a paper. Every PDF link is checked against a small set of rules (rejecting API-shaped hostnames, metadata paths, and telltale query parameters) before it's ever shown as “View free PDF.” When a link doesn't pass, Cortexa falls back to linking the paper's landing page instead of showing a broken download.
The same paper often exists in more than one source — an arXiv preprint later published in IEEE Access, for instance. Cortexa keeps one canonical entry per paper (matched by DOI) rather than showing duplicates, and notes the others as “also available via” on that entry instead of dropping the information entirely.
Cortexa doesn't host, mirror, or redistribute any paper's file. Every download link points to the publisher's or repository's own copy. Nothing here is scraped from behind a paywall.