Extension Path: For Researchers¶
📌 What this path helps you do¶
This page does not make AI your researcher. It helps with one simpler task: find sources, understand them, and confirm that answers are actually supported by evidence.
- If you use a terminal or Python, come after Track A A3 or Track B Stage 7.
- If you do not code, start with the first exercise below. You only need a browser and one public paper.
🎯 Learning goals¶
After this page, you can: 1. Separate what AI says from what the original text says. 2. Check each numbered reference instead of trusting an answer just because it has reference numbers. 3. Know which data may be uploaded and which requires permission from an institution or data owner. 4. Keep enough records for yourself or a colleague to reproduce the work.
🧩 Eight core terms¶
- Source: original material used for verification, such as a paper, dataset, or research record.
- Claim: a checkable statement, such as “method A performs better on dataset B.”
- Citation: a signpost back to a source location; it does not guarantee that the source supports the claim.
- Source Verification: open the original and check that content, scope, and limitations match the answer.
- Literature RAG: retrieve passages from permitted literature, then give them to a model to answer.
- Reproducibility: others can rerun comparable results from your data, steps, versions, and settings.
- Private Data: content that cannot be freely published or uploaded, such as participant data, medical records, unpublished manuscripts, or company secrets.
- Human Review: a person is responsible for claims, citations, code, tables, and the final decision; AI cannot sign or assume responsibility.
🛠 First exercise: verify three answers about one paper¶
Before uploading, confirm that the license or copyright and the tool's terms allow it. Publicly readable is not permission to upload a paper to another service.
Use the public paper Attention Is All You Need. Add the paper to a citation-capable tool and paste:
Answer only from this paper. Attach a citation to each answer; if evidence is missing, write “unsupported” and do not guess.
1. What problem does the paper aim to solve?
2. What are the main parts of the proposed method?
3. Which experiments support the result, and what limitations do the authors state?
After answering, list the original text for each citation. Do not present your inference as an author claim.
📚 Choose an entry point¶
| What you want | Start with | Why | Rating |
|---|---|---|---|
| Ask about one paper in a browser | Gemini Notebook (formerly NotebookLM) | Return from source uploads to citations | ⭐⭐⭐⭐⭐ |
| Organize your literature library | Zotero | Organize PDFs, authors, years, and notes first | ⭐⭐⭐⭐⭐ |
| Build rerunnable literature RAG in Python | PaperQA2 | Science-document and citation-centered workflow | ⭐⭐⭐⭐⭐ |
Gemini Notebook is Google’s current name for NotebookLM as of 2026-07-16; the old name remains for recognition. A citation is an entry point for checking, not a guarantee.
📖 Required reading¶
Read in order. The first two prevent treating citations as guarantees; the next four help preserve sources, code, data, and results: 1. Gemini Notebook citation help: open citations and read context. 2. Gemini Notebook privacy and terms: understand data handling before upload. 3. Zotero quick start: organize authors, years, PDFs, and notes. 4. PaperQA2 README: connect literature RAG answers to documents. 5. DVC command reference: version data and rerunnable pipelines with Git. 6. Zenodo quickstart: preserve publishable data, code, or materials in a citable version.
⭐ Curated research tools and projects¶
Tool names, licenses, and repository status were checked against official pages and the GitHub API on 2026-08-29 UTC. Ratings are editorial ratings for this map, not GitHub stars or rankings.
| Category | Official tool / project | Good for | Status / license | Know this limitation | Rating |
|---|---|---|---|---|---|
| Start and organize | Gemini Notebook (formerly NotebookLM) | Source-grounded Q&A and citations | Available; cloud service | Check every citation; review policy before private data | ⭐⭐⭐⭐⭐ |
| Zotero | Manage PDFs, metadata, notes, and citations | Available; desktop / web | Manages sources; does not judge research quality | ⭐⭐⭐⭐⭐ | |
| Future-House/paper-qa | Build citation-grounded literature RAG in Python | Active; Apache-2.0 | Configure model and sources; evaluate quality yourself | ⭐⭐⭐⭐⭐ | |
| Explore and write | assafelovic/gpt-researcher | Multi-source search and research briefs | Active; Apache-2.0 | Find candidate sources; not the final citation judge | ⭐⭐⭐⭐ |
| stanford-oval/storm | Organize viewpoints, outlines, and long-form writing | Usable; MIT; slower updates | Check dependencies and sources before use | ⭐⭐⭐⭐ | |
| kaixindelele/ChatPaper | Chinese paper summaries, translation, and writing support | Usable; CC BY-NC-ND 4.0 | Noncommercial, no-derivatives license; not a general open-source license | ⭐⭐⭐⭐⭐ | |
| MuiseDestiny/zotero-gpt | Interact with literature in Zotero | Usable; AGPL-3.0 | Maintain plugin and model settings separately | ⭐⭐⭐⭐ | |
| Reproducibility and evidence | asreview/asreview | Active-learning support for systematic-review screening | Active; Apache-2.0 | Ranking saves time; human screening still decides inclusion and keeps the record | ⭐⭐⭐⭐ |
| treeverse/dvc | Keep data versions, models, and pipelines rerunnable | Active; Apache-2.0 | Needs Git and storage; versions do not prove conclusions | ⭐⭐⭐⭐⭐ | |
| mlflow/mlflow | Track parameters, metrics, data, and artifacts across runs | Active; Apache-2.0 | Tracking does not make an experiment valid; keep secrets and participant data out | ⭐⭐⭐⭐⭐ | |
| Zenodo | Publish data, code, and materials with a DOI | Available; cloud service | Metadata is public; de-identify private data under institutional rules | ⭐⭐⭐⭐⭐ | |
| jupyterhub/repo2docker | Rebuild a runnable environment from repository settings | Active; BSD-3-Clause | A container preserves the environment; also preserve data, hardware needs, and external services | ⭐⭐⭐⭐ | |
| Research automation | flonat/flonat-research | Research skills, agents, hooks, and LaTeX workflows | Active; MIT | Infrastructure example, not universal for every field | ⭐⭐⭐ |
| SakanaAI/AI-Scientist-v2 | End-to-end multi-agent research experiments | Research reference; custom source-code license | License requires disclosure of machine-generated manuscripts; authors remain responsible | ⭐⭐⭐⭐ | |
| History | langchain-ai/open_deep_research | Study early deep-research agent architecture | Archived; MIT; historical reference | Not a current default | ⭐⭐⭐ |
✅ Completion check and next stop¶
- I checked three answers, not just citation numbers.
- I found an example supported by the original or marked unsupported.
- I did not upload unapproved Private Data.
- I saved sources, questions, tool name, date, and my judgment.
Next: use Stage 6 for literature RAG; Stage 7 for multiple agents; and the MCP / Skills catalog for external tools.
⏱ Expand: time, accounts, cost, and data safety
The first exercise takes about 20–40 minutes. For Private Data, pause and confirm IRB, institutional policy, contracts, data-owner consent, and tool terms. Gemini Notebook privacy guidance says general content is not directly used to train foundation models unless feedback is provided, and feedback may be reviewed by people; this does not automatically approve research uploads. Plans, quotas, and account rules change, so check official pages rather than preserving fixed prices.
🧪 Expand: turn one-paper practice into a rerunnable workflow
Literature inbox¶
Save DOI, URL, authors, year, and acquisition date; let tools summarize while linking each claim to the original; humans decide read, exclude, or verify and record why.
Cross-paper synthesis¶
Ask what each paper says before comparing agreement, conflict, and conditions. Do not ask for a complete story before finding citations.
Code and experiments¶
Save data versions, environment, seed, prompt, model/tool versions, outputs, and human edits. Rerunning does not prove correctness, but missing records hide errors.
Before submission¶
Check every claim, citation, table, figure, program, and journal rule. Authors make the final judgment and disclose use under journal policy.
🧯 Expand: common errors, alternatives, and troubleshooting
| Problem | What to do first |
|---|---|
| Citation does not support the answer | Mark unsupported, narrow the question, and do not force a related citation |
| Tool cannot read a scanned PDF | OCR first, then spot-check pages and formulas |
| Conclusions from papers are mixed | Require paper name and page/paragraph for each claim before synthesis |
| Data cannot go to the cloud | Use an institutional environment; consider the local RAG route in Stage 6 |
| Automation is too complex | Return to one paper, three questions, and one-by-one checking |
No tool replaces IRB, data governance, author responsibility, or domain expertise.