The Researcher's Toolkit
Literature search, Zotero, the comparison matrix, LaTeX templates, BibTeX, citation hygiene
Appendix B — The Researcher’s Toolkit
“For which of you, desiring to build a tower, does not first sit down and count the cost, whether he has enough to complete it?” — Luke 14:28 (ESV)
You don’t show up to the meet without your gear. A sprinter checks the blocks; a climber inspects the rope. This appendix is your kit bag for the semester — the literature stack you’ll search in, the manager that keeps your sources straight, the matrix that turns a pile of PDFs into an argument, the templates your paper will live inside, and the discipline that keeps a fabricated citation from torpedoing your name. Read it once now. Come back to it every week.
One standing rule before we open the bag. Tool pricing and corpus sizes drift fast. Every number below is as of 2026, and several of these vendors revise their plans and their index counts more than once a year. When a free tier or a corpus size is load-bearing for your decision, open the vendor’s pricing page and confirm it live. A confident wrong number is worse than no number.
B.1 — The Literature Stack
Think of literature work in four layers — find, understand, snowball, summarize — plus a fifth you’ll meet in B.2: manage. No single tool does all five well. You assemble a stack.
Layer 1 — Find (broad search and authoritative metadata)
| Tool | Cost (as of 2026) | What it’s for |
|---|---|---|
| Google Scholar | Free | The broadest net. Largest reach of anything here (an external estimate puts it north of 400M documents — not an official figure). Start wide here, then narrow. |
| ACM Digital Library | Free (Basic) as of Jan 1, 2026 | The ACM corpus went fully open access on January 1, 2026. Free Basic gives you the full corpus with limited features; paid Premium adds analytics and discovery tooling. |
| IEEE Xplore | Subscription / institutional | Did not go open. You’ll most likely reach it through your university’s subscription. Plan for a campus login or a library proxy. |
| arXiv | Free | Preprints in the cs.* categories. Moderated, not peer-reviewed — treat an arXiv-only paper as a claim, not a settled result. |
| dblp | Free (CC0) | The authoritative CS bibliography, maintained by Schloss Dagstuhl. Mid-2026 it lists roughly 8.6M publications (~8,599,343 pubs / ~4,127,107 authors on the homepage counter). Best place to confirm clean metadata; it labels arXiv items as “informal publications.” |
Coach’s Note — The ACM going open access on January 1, 2026 changed the default move for a lot of you. A year ago “I can’t read it, it’s behind ACM’s paywall” was a real excuse. It isn’t anymore. Get the paper.
Layer 2 — Understand (summaries and citation context)
| Tool | Cost (as of 2026) | What it’s for |
|---|---|---|
| Semantic Scholar | Free | Allen Institute for AI (AI2). Roughly 214M+ papers and ~2.49B citations (approximate, growing). TLDR auto-summaries, Semantic Reader, and SPECTER2 embeddings. The S2AG API is free — handy for programmatic pulls. |
Layer 3 — Snowball (references backward, citers forward)
| Tool | Cost (as of 2026) | What it’s for |
|---|---|---|
| Connected Papers | Freemium — ~5 graphs/month free | A similarity graph around a seed paper. Great for seeing a neighborhood at a glance. |
| Research Rabbit | Freemium (free tier caps each search at ~50 seed papers) | Now owned by Litmaps (acquisition announced May 8, 2025; freemium tier landed Nov 2025). Living collections that grow as you add papers. |
| dblp | Free | Also your backward/forward workhorse for verified metadata while snowballing. |
| scite | Paid (see Layer 4) | Forward citation tracking with stance — who cited this, and were they supporting or contrasting it. |
Layer 4 — Summarize (grounded, retrieval-based AI)
These ground their answers in a real index, so they hallucinate less than a raw chatbot. Less is not zero — B.5 is non-negotiable.
| Tool | Cost (as of 2026) | What it’s for |
|---|---|---|
| Elicit | Free tier now unlimited; paid Plus/Pro/Scale tiers add data extraction and systematic-review tooling | ~138M papers. Structured field extraction into custom columns — feeds your comparison matrix directly. |
| Consensus | Free (capped ~20 searches/month); Pro and Deep tiers above that | ~200M+ papers, the Consensus Meter, and Study Snapshots that auto-extract methods/outcomes/sample sizes. |
| scite | Paid | Smart Citations classify each citation as supporting / contrasting / mentioning — the single best tool for “was this result actually replicated, or just name-dropped?” |
Which are free
- Free, full stop (as of 2026): Google Scholar, ACM Digital Library (Basic, fully open since Jan 1, 2026), arXiv, dblp (CC0), Semantic Scholar + the S2AG API.
- Free tier you’ll bump into limits on: Connected Papers (~5 graphs/mo), Research Rabbit (~50 seeds/search), Elicit (unlimited free tier, paid for extraction at scale), Consensus (~20 searches/mo free).
- Subscription / institutional (budget for it): IEEE Xplore, ACM Premium features, scite, and the upper Elicit/Consensus tiers.
Coach’s Note — Build the habit of starting free and authoritative: Scholar to find, dblp to confirm the citation is real and clean, Semantic Scholar to understand it fast. Reach for a paid tool only when the free stack genuinely runs out — and most weeks it won’t.
B.2 — The Zotero Workflow (Layer 5: Manage)
Your reference manager is Zotero — free, open-source, the standard. One fact to set expectations: Zotero has no native AI in 2026. Any “ask my library” behavior comes from third-party plugins (PapersGPT, Aria, Beaver), not the core app. It also caps free cloud storage at 300MB. (The Elsevier-owned alternative, Mendeley, gives 2GB free and ships built-in AI — but you trade open-source neutrality and portability for it. We default to Zotero for the values and the lock-in math; that’s a debate worth having in class, not a decision to make on autopilot.)
The workflow, in order:
- Install Zotero + the browser connector. The connector saves a paper, its metadata, and the PDF in one click from Scholar, the ACM DL, IEEE Xplore, arXiv, or a publisher page.
- Make a collection per project. One collection for your practicum. Sub-collections by theme if it gets large. Tag aggressively —
baseline,dataset,must-cite,contested. - Capture, then immediately verify the metadata. Connectors fumble author lists, venues, and years constantly — especially on arXiv items. Fix the entry against dblp the moment you save it. A clean library now is a clean bibliography later.
- Read inside Zotero. Annotate the PDF; pull your highlights into a per-paper note. Those notes become matrix cells in B.3.
- Cite with Better BibTeX. Install the Better BibTeX plugin, set a stable citation-key scheme (e.g.
authorYearword), and export your collection to a.bibfile for LaTeX. Turn on keep updated / auto-export so the.bibtracks the library as you add sources. - Sync and back up. Zotero sync for metadata; watch the 300MB ceiling on attachments (use linked files or institutional WebDAV/storage if you blow past it).
Coach’s Note — The single most common own-goal I see: students let the connector’s sloppy metadata ride for three months, then generate a bibliography full of wrong years and mangled author names the night before the deadline. Fix every entry on capture. It costs thirty seconds then and an evening of misery later.
B.3 — The Comparison-Matrix Template
Synthesis is not summary. A summary tells me what each paper said. A matrix lets me see what the papers disagree about — and that disagreement is where your gap lives. Rows are papers; columns are the dimensions you compare across them.
| Paper (cite key) | Problem | Method | Dataset / Benchmark | Baselines | Metrics | Claimed result | Limitation / threat |
|---|---|---|---|---|---|---|---|
smith2025rag | … | … | … | … | … | … | … |
lee2026agentic | … | … | … | … | … | … | … |
… | … | … | … | … | … | … | … |
Build it like this:
- One row per key paper. Start with 6–10 of your strongest sources.
- Let a grounded tool fill the easy columns. Elicit’s custom-column extraction or Consensus Study Snapshots can populate method, dataset, metrics, sample size across all rows at once. This is a legitimate accelerator.
- Hand-verify every single cell against the PDF. This is the rule, not a suggestion — auto-extraction misattributes and over-claims. If a cell can’t be confirmed in the primary source, it doesn’t go in the matrix.
- Read the matrix down the columns, not across the rows. Where do the limitations cluster? Which benchmark does everyone lean on, and is it contamination-prone? That column-wise read is the synthesis.
- Run your contested claims through scite to reclassify them supporting / contrasting / mentioning before you write “prior work shows.”
Coach’s Note — One caution on scope. As of arXiv’s October 31, 2025 policy, a CS review/survey/position paper must already be peer-review-accepted at a journal or conference (workshop review is explicitly not enough — journal ref + DOI required). So your practicum lit review, however good, is not something you can post to arXiv as a standalone survey. It’s the engine of your paper, not a paper by itself.
B.4 — Paper Templates: acmart, IEEEtran, and BibTeX
Your venue dictates your class file. Both of these are pre-loaded on Overleaf, so you don’t have to fight a local TeX install on day one.
ACM — acmart
- Class file
acmart.cls, v2.18 (dated 2026/06/01 as of this writing — confirm the literal version string against the ACM source before you submit, as it revises often). - Submit single-column for review:
Add\documentclass[manuscript,review]{acmart}anonymousfor double-blind venues — your name renders as “ANONYMOUS AUTHOR(S)”. - Camera-ready switches to two-column via the proceedings option (most CS venues use
sigconf):\documentclass[sigconf]{acmart} - ACM’s TAPS pipeline processes the source and emits both a two-column PDF and responsive HTML5. Page length is counted in the one-column submission format, so don’t be fooled by how short the two-column version looks.
IEEE — IEEEtran
- Class file
IEEEtran.cls, v1.8b (maintainer Michael Shell; the core has been unchanged on CTAN since 2015-08-28 — stable, unlike acmart). - Conference papers, two columns:
\documentclass[conference]{IEEEtran}
BibTeX — wiring it together
Both classes consume a BibTeX .bib file — the one you auto-export from Zotero via Better BibTeX (B.2). Typical end of document:
\bibliographystyle{ACM-Reference-Format} % for acmart; use IEEEtran for IEEE
\bibliography{refs} % refs.bib exported from Zotero
Cite with \cite{smith2025rag}, using the stable keys you set in Better BibTeX. Keep the .bib regenerating from Zotero so a metadata fix in your library propagates to the paper automatically — never hand-edit the .bib for content, fix it in Zotero and re-export.
Coach’s Note — Pick
acmartvsIEEEtranby your target venue, and pick the venue before you write a line of LaTeX. The skeleton you build is the venue’s skeleton. Building the wrong one and converting later is a tax you pay for not deciding early.
B.5 — Citation Hygiene (Non-Negotiable)
This is the section that protects your name. Read it twice.
LLMs fabricate citations at a rate that will end careers if you trust them blind. The hard numbers from Walters & Wilder 2023 (Scientific Reports): 55% of GPT-3.5 citations and 18% of GPT-4 citations were entirely fabricated — papers that do not exist. Among the citations that were real, 43% (GPT-3.5) / 24% (GPT-4) still had substantive errors. A separate study (Bhattacharyya 2023, Cureus) found 87% of citations to real works had at least one metadata error. Those percentages are from 2023-era models and are illustrative point-in-time figures — newer retrieval-grounded “deep research” agents shift them downward — but the failure mode has not gone away. Grounded/RAG tools (Elicit, Consensus, scite) hallucinate less. Less is not zero.
So the rule is absolute:
Verify every reference against a real index before it enters your .bib. Every one. Cross-check the title, authors, venue, and year in dblp (authoritative CS metadata) and Semantic Scholar (or the publisher’s own page). If a reference an AI handed you does not resolve in dblp or Semantic Scholar to the exact paper described, it does not exist — delete it.
Why this is existential and not pedantic:
- Top venues reject papers with non-existent citations without review — ICCV does exactly this.
- arXiv can impose a one-year submission ban for “incontrovertible evidence” of unchecked LLM content — hallucinated references are the canonical example (a moderator policy reported in May 2026; re-verify exact terms, but don’t bet your record on its being unenforceable).
- “The model did it” is never a defense. At every major venue — ACM, IEEE, NeurIPS, ICLR, ICML, ACL, CVPR/ICCV, arXiv — the human author is fully accountable for every citation, figure, and line. A tool cannot bear that accountability, which is exactly why an LLM can never be an author.
A workflow that keeps you safe:
- Never paste an AI-generated citation straight into your
.bib. - For each reference: find it in dblp → confirm title/authors/venue/year → grab the clean metadata (or the dblp/publisher BibTeX) → import into Zotero → re-export.
- Spot-check the claim, not just the existence. A real paper cited for something it never says is still a fabrication of the argument. Open the PDF; confirm it supports the sentence you attached it to.
- Disclose substantive generative use where your venue requires it (typically the Acknowledgements; NeurIPS uses the experimental-setup section). Spell-check and grammar polish don’t need disclosing; generated text, ideas, code, or citations do.
Coach’s Note — Here’s the whole semester in one line: the human stays in the loop where the judgment lives. The AI can accelerate your search, draft your matrix columns, and polish your prose. It cannot decide whether a citation is real, and it cannot be accountable when it isn’t. You verify. You decide. You sign your name. That’s not a constraint on the work — that’s what makes it research.
Drill yourself once and you’ll have the reflex for life: ask any chatbot for five references on a niche topic, then run all five through dblp and Semantic Scholar. Count the ones that don’t exist. The number will scare you straight.
That’s the kit. Find with the free authoritative stack, manage in Zotero, synthesize in the matrix, write in the right template, and verify every citation against a real index like your name depends on it — because it does.
See you in the lab.