Skip to content

Extension Path: For Researchers

← Back to the main route

📌 What this path helps you do

This page does not make AI your researcher. It helps with one simpler task: find sources, understand them, and confirm that answers are actually supported by evidence.

  • If you use a terminal or Python, come after Track A A3 or Track B Stage 7.
  • If you do not code, start with the first exercise below. You only need a browser and one public paper.

🎯 Learning goals

After this page, you can: 1. Separate what AI says from what the original text says. 2. Check each numbered reference instead of trusting an answer just because it has reference numbers. 3. Know which data may be uploaded and which requires permission from an institution or data owner. 4. Keep enough records for yourself or a colleague to reproduce the work.

🧩 Eight core terms

  • Source: original material used for verification, such as a paper, dataset, or research record.
  • Claim: a checkable statement, such as “method A performs better on dataset B.”
  • Citation: a signpost back to a source location; it does not guarantee that the source supports the claim.
  • Source Verification: open the original and check that content, scope, and limitations match the answer.
  • Literature RAG: retrieve passages from permitted literature, then give them to a model to answer.
  • Reproducibility: others can rerun comparable results from your data, steps, versions, and settings.
  • Private Data: content that cannot be freely published or uploaded, such as participant data, medical records, unpublished manuscripts, or company secrets.
  • Human Review: a person is responsible for claims, citations, code, tables, and the final decision; AI cannot sign or assume responsibility.

🛠 First exercise: verify three answers about one paper

Before uploading, confirm that the license or copyright and the tool's terms allow it. Publicly readable is not permission to upload a paper to another service.

Use the public paper Attention Is All You Need. Add the paper to a citation-capable tool and paste:

Answer only from this paper. Attach a citation to each answer; if evidence is missing, write “unsupported” and do not guess.
1. What problem does the paper aim to solve?
2. What are the main parts of the proposed method?
3. Which experiments support the result, and what limitations do the authors state?
After answering, list the original text for each citation. Do not present your inference as an author claim.
Then open each citation, read answer and original text together, and mark unsupported sentences unsupported instead of adding an unrelated citation.

📚 Choose an entry point

What you want Start with Why Rating
Ask about one paper in a browser Gemini Notebook (formerly NotebookLM) Return from source uploads to citations ⭐⭐⭐⭐⭐
Organize your literature library Zotero Organize PDFs, authors, years, and notes first ⭐⭐⭐⭐⭐
Build rerunnable literature RAG in Python PaperQA2 Science-document and citation-centered workflow ⭐⭐⭐⭐⭐

Gemini Notebook is Google’s current name for NotebookLM as of 2026-07-16; the old name remains for recognition. A citation is an entry point for checking, not a guarantee.

📖 Required reading

Read in order. The first two prevent treating citations as guarantees; the next four help preserve sources, code, data, and results: 1. Gemini Notebook citation help: open citations and read context. 2. Gemini Notebook privacy and terms: understand data handling before upload. 3. Zotero quick start: organize authors, years, PDFs, and notes. 4. PaperQA2 README: connect literature RAG answers to documents. 5. DVC command reference: version data and rerunnable pipelines with Git. 6. Zenodo quickstart: preserve publishable data, code, or materials in a citable version.

⭐ Curated research tools and projects

Tool names, licenses, and repository status were checked against official pages and the GitHub API on 2026-08-29 UTC. Ratings are editorial ratings for this map, not GitHub stars or rankings.

CategoryOfficial tool / projectGood forStatus / licenseKnow this limitationRating
Start and organizeGemini Notebook (formerly NotebookLM)Source-grounded Q&A and citationsAvailable; cloud serviceCheck every citation; review policy before private data⭐⭐⭐⭐⭐
ZoteroManage PDFs, metadata, notes, and citationsAvailable; desktop / webManages sources; does not judge research quality⭐⭐⭐⭐⭐
Future-House/paper-qaBuild citation-grounded literature RAG in PythonActive; Apache-2.0Configure model and sources; evaluate quality yourself⭐⭐⭐⭐⭐
Explore and writeassafelovic/gpt-researcherMulti-source search and research briefsActive; Apache-2.0Find candidate sources; not the final citation judge⭐⭐⭐⭐
stanford-oval/stormOrganize viewpoints, outlines, and long-form writingUsable; MIT; slower updatesCheck dependencies and sources before use⭐⭐⭐⭐
kaixindelele/ChatPaperChinese paper summaries, translation, and writing supportUsable; CC BY-NC-ND 4.0Noncommercial, no-derivatives license; not a general open-source license⭐⭐⭐⭐⭐
MuiseDestiny/zotero-gptInteract with literature in ZoteroUsable; AGPL-3.0Maintain plugin and model settings separately⭐⭐⭐⭐
Reproducibility and evidenceasreview/asreviewActive-learning support for systematic-review screeningActive; Apache-2.0Ranking saves time; human screening still decides inclusion and keeps the record⭐⭐⭐⭐
treeverse/dvcKeep data versions, models, and pipelines rerunnableActive; Apache-2.0Needs Git and storage; versions do not prove conclusions⭐⭐⭐⭐⭐
mlflow/mlflowTrack parameters, metrics, data, and artifacts across runsActive; Apache-2.0Tracking does not make an experiment valid; keep secrets and participant data out⭐⭐⭐⭐⭐
ZenodoPublish data, code, and materials with a DOIAvailable; cloud serviceMetadata is public; de-identify private data under institutional rules⭐⭐⭐⭐⭐
jupyterhub/repo2dockerRebuild a runnable environment from repository settingsActive; BSD-3-ClauseA container preserves the environment; also preserve data, hardware needs, and external services⭐⭐⭐⭐
Research automationflonat/flonat-researchResearch skills, agents, hooks, and LaTeX workflowsActive; MITInfrastructure example, not universal for every field⭐⭐⭐
SakanaAI/AI-Scientist-v2End-to-end multi-agent research experimentsResearch reference; custom source-code licenseLicense requires disclosure of machine-generated manuscripts; authors remain responsible⭐⭐⭐⭐
Historylangchain-ai/open_deep_researchStudy early deep-research agent architectureArchived; MIT; historical referenceNot a current default⭐⭐⭐

✅ Completion check and next stop

  • I checked three answers, not just citation numbers.
  • I found an example supported by the original or marked unsupported.
  • I did not upload unapproved Private Data.
  • I saved sources, questions, tool name, date, and my judgment.

Next: use Stage 6 for literature RAG; Stage 7 for multiple agents; and the MCP / Skills catalog for external tools.

⏱ Expand: time, accounts, cost, and data safety

The first exercise takes about 20–40 minutes. For Private Data, pause and confirm IRB, institutional policy, contracts, data-owner consent, and tool terms. Gemini Notebook privacy guidance says general content is not directly used to train foundation models unless feedback is provided, and feedback may be reviewed by people; this does not automatically approve research uploads. Plans, quotas, and account rules change, so check official pages rather than preserving fixed prices.

🧪 Expand: turn one-paper practice into a rerunnable workflow

Literature inbox

Save DOI, URL, authors, year, and acquisition date; let tools summarize while linking each claim to the original; humans decide read, exclude, or verify and record why.

Cross-paper synthesis

Ask what each paper says before comparing agreement, conflict, and conditions. Do not ask for a complete story before finding citations.

Code and experiments

Save data versions, environment, seed, prompt, model/tool versions, outputs, and human edits. Rerunning does not prove correctness, but missing records hide errors.

Before submission

Check every claim, citation, table, figure, program, and journal rule. Authors make the final judgment and disclose use under journal policy.

🧯 Expand: common errors, alternatives, and troubleshooting
Problem What to do first
Citation does not support the answer Mark unsupported, narrow the question, and do not force a related citation
Tool cannot read a scanned PDF OCR first, then spot-check pages and formulas
Conclusions from papers are mixed Require paper name and page/paragraph for each claim before synthesis
Data cannot go to the cloud Use an institutional environment; consider the local RAG route in Stage 6
Automation is too complex Return to one paper, three questions, and one-by-one checking

No tool replaces IRB, data governance, author responsibility, or domain expertise.