Academic Research Skills (ARS) is a source-available academic research copilot framework for noncommercial scholarly use. The reference distribution is a suite of Claude Code skills that assists human researchers through the full research-to-publication pipeline. Sibling distributions for other agent platforms (e.g. Codex) follow the same workflow content, the same human-in-the-loop design philosophy, and the same license terms; see CONTRIBUTING.md § Platform ports.
It is licensed under CC BY-NC 4.0. This is not an open source license — it restricts commercial use by design, to keep the tool free for academic communities.
ARS is not an autonomous paper-writing system. It is not a replacement for the researcher. It does not claim authorship, and its outputs are not submission-ready without human review.
These are not "out of scope" footnotes. They are the load-bearing boundary that defines what ARS does NOT do, and would not do even if a future system made them feasible. Most are autonomous mechanisms catalogued by Kong et al. (2026), AI for Auto-Research: Roadmap & User Guide (arXiv:2605.18661), and rejected against the human-led positioning above. The recorded review test for the autonomous-research mechanisms — "who controls the next research-state transition?" — lives in the L1 design lesson.
- End-to-end autonomous research pipeline (Kong §7.4.8). A system that carries a project from question to manuscript without scholar confirmation at each state transition. Rejected: the scholar would become a reviewer of AI output, not the author. The pipeline's mandatory checkpoints exist precisely to prevent this.
- Autonomous idea-generation agent (Kong §3.1). An agent that proposes
research hypotheses or questions for the scholar without an explicit
authorship-boundary transition. Rejected — and distinct from the shipped
wording-pattern advisory (#257): while non-generation Socratic mode is active,
ARS may flag surface wording/framing patterns, summarize only directions the
scholar has already expressed, and ask follow-up questions, but it must not
propose, substitute, rank, expand, or select research hypotheses or questions
for the scholar. Non-convergence is never consent. If the scholar explicitly
asks the system itself to propose candidates, ARS must visibly leave that mode
by emitting
[SOCRATIC-NON-GENERATION-EXIT: explicit_user_request]before any candidate content and label the result AI-generated; this is a disclosed mode change, not a hidden Socratic fallback. The boundary is recorded in the L2 design lesson. - Paper2X auto-generation (Kong §6). Autonomous generation of slides / posters / video from a manuscript. Rejected — and distinct from a fidelity audit: ARS may audit an already-authored or externally generated dissemination artifact against the manuscript for fidelity, but it must not transform a manuscript into a dissemination artifact by choosing the content, narrative, layout, or output medium itself. (Dissemination design is handled by separate, non-ARS skill chains; the fidelity-audit suggestion itself is out of this repo's scope.)
- Autonomous experiment execution / coding (Kong §3.3). An LLM that runs experiments or code without scholar oversight. Rejected — and distinct from the shipped Experiment Provenance Intake (#260): ARS may ingest scholar-declared external experiment provenance and check manuscript claims against the declared results, but it must not initiate, run, modify, iterate, or treat tool-executed experiment / code outputs as evidence inside the pipeline.
- Physical wet-lab automation API (Kong §7.4.6). An interface that drives liquid handlers or automated labs. Rejected: even with safeguards, this extends beyond a research copilot's scope into laboratory infrastructure, and conflicts with the copilot-not-pilot positioning.
- Simulated human-subjects review committee. LLM lenses named after statutory committee seats, pre-committing a protocol risk level and combining seat judgments into a committee-like result. Rejected: statutory composition rules create an independent, representative, conflict-accountable human body; they are not an epistemic recipe whose legitimacy transfers to model personas (45 CFR 46.107; Taiwan Human Subjects Research Act, Art. 7). A risk level cannot be meaningfully pre-committed before protocol facts are seen, determination letters are not ethical ground truth, and the unresolved reviewer severity-band error (#648) is especially consequential when risk language is the output. The ownership boundary is categorical: AI may generate questions or advisory observations, but a judgment that binds an absent person requires an accountable human owner. If this topic returns, the defensible object is an RFC and held-out evaluation of multi-lens question generation—concern recall, false reassurance, and abstention—not risk levels, committee verdicts, or a system called a committee.
These are first-party scope boundaries and review criteria for future changes, not runtime guarantees. First-party ARS treats each as out of scope; adding one would require changing this recorded boundary, not merely adding a feature.
Unlike the Rejected mechanisms above — capabilities ARS refuses on principle — these are lifecycle stages and state layers ARS deliberately does not enter. They were adjudicated out of scope in the 2026-06-10 researcher-blindspot audit and are recorded here so the boundary is reviewable, not improvised (the same recording discipline as the Rejected mechanisms; boundary + review criterion, not a runtime guarantee).
- Post-publication lifecycle. Tracking citation contexts of the scholar's own published papers, errata/corrigenda workflows, and OA self-archiving compliance are out of scope. ARS's front is research-to-publication; what happens to a paper after it ships belongs to the scholar and their institutional tooling. The existing
monitoring_agentis unaffected — it alerts on developments in the cited literature (an input to current work), not on the scholar's own published output. Review criterion: a proposed feature whose value begins after the manuscript is accepted extends the front, and requires changing this recorded boundary first. - Research-program-level state. ARS keeps no memory across papers: no registry of the scholar's prior claims, no carried-forward limitations list, no reviewer-history profile. The per-paper Material Passport remains the only state carrier, and every run starts from what the scholar explicitly feeds it. This is a deliberate consequence of the anti-leakage philosophy — gates that trusted an ambient cross-paper memory would be evaluating state nobody declared this run. The supported way for a returning author to carry their own prior work forward without any new mechanism is the Cross-paper workflow guide. Review criterion: a proposed feature that reads or writes scholar state outside the current run's passport crosses this boundary.
- Institutional / journal format-profile content. Unlike the two above, ARS does ship the mechanism — the scholar-declared layout
format_profile(#439), so a user can bind a thesis or journal template without forking. What ARS deliberately does NOT ship is any specific institution's or journal's profile content: the repo carries the schema and a synthetic example only, never a real school's font/spacing/caption rules. Binding the suite to one institution's template is the boundary the #439 design keeps out. Review criterion: a PR that adds a real institution's or journal'sformat_profile.yaml(or hardcodes its rules into an agent) to this repo crosses this boundary — profiles stay user-supplied and out-of-tree.
ARS assesses the manuscript and the reported process. Its checks cover citation existence, claim–source alignment, assessment of the reported methodology, alignment between manuscript claims and scholar-declared experiment results, figure/table fidelity, reporting-guideline coverage, and process/package conformance. These are checks with explicit coverage limits—not guarantees: some are sampled, some depend on external-index coverage, and some are LLM-mediated judgments.
ARS does not validate actual execution. It cannot establish that the reported procedures were performed, that raw data are authentic or complete, that analyses reproduce from the underlying materials, or that a real-world intervention occurred as described. Experiment Provenance Intake records scholar declarations and checks their internal alignment; it does not execute or independently validate the experiment. The honest failure shape is stark: a study built on fabricated data can pass every ARS gate if the fabrication is consistently reported, cited, and packaged. Researchers and readers must therefore treat ARS outputs as bounded manuscript/process checks and retain human, institutional, and reproducibility review for the underlying empirical work.
- Research assistance: literature search, source verification, citation checking
- Teaching: demonstrating research methodology, peer review processes, academic writing standards
- Method training: using Socratic modes to develop research question formulation and argumentation skills
- Noncommercial academic collaboration: research groups, labs, departments using the tool for shared workflows
- Submitting AI-generated papers as solely human-authored without disclosing AI assistance
- Using the tool to produce papers without engaging with the content (the pipeline has mandatory checkpoints specifically to prevent this)
- Treating AI-generated review feedback as a substitute for actual peer review
- Commercial SaaS or hosted services built on ARS
- Consulting or freelance services that package ARS as a paid product
- Enterprise or institutional paid deployments without separate licensing
- Commercial API wrappers or resale of ARS functionality
These reflect our policy intent. See the CC BY-NC 4.0 license for the precise legal terms. For commercial licensing inquiries, contact the maintainer.
Assistive, not deceptive. ARS helps you write better, not hide that you used AI.
- Style Calibration learns your voice from past papers — so the output sounds like you, not like a machine
- Writing Quality Check catches AI-typical patterns — to improve prose quality, not evade detection
- Disclosure Mode generates venue-specific or policy-anchor AI usage statements — because transparency is the standard
Human-in-the-loop, always. The pipeline's checkpoint system is mandatory by design:
- FULL checkpoints present all deliverables and require explicit user confirmation
- MANDATORY checkpoints at integrity gates and review decisions cannot be skipped
- "Full mode" means full-pipeline execution, not full autonomy — the human decides at every gate
- Max 2 revision loops, after which remaining issues become "Acknowledged Limitations" rather than being silently resolved
Failure modes are made visible, not hidden. The 7-mode AI Research Failure Mode Checklist (v3.2) and Reviewer Calibration Mode exist so that users can see where the AI might be wrong — not so that the AI can claim it's always right. The v3.7.3 + v3.8 L3 claim-faithfulness gate adds per-citation locator anchors and an opt-in audit pass that verifies whether each cited source actually supports the claim made of it.
Boundaries are recorded, not improvised. When adopting a capability from a published system would touch a load-bearing boundary — who ranks, what propagates, who writes state — the decision of whether and how to adopt it is written down as a design-lesson doc, so the same boundary is applied consistently later. The Co-Scientist (Gottweis et al. 2026) analysis is recorded in four such docs: hidden-ranking vs. advisory ranking (L1), unapproved feedback propagation (L2), which mechanisms transfer to ARS and which do not (L3), and control-plane ownership — who may write, rank, or route (L4). The Kong (2026) auto-research analysis adds two: copilot vs. auto-research as a research-state-authority line (L1) and advisory-on-wording vs. idea-generation (L2); the autonomous mechanisms they reject are enumerated in Rejected mechanisms above.
If you use ARS in your research, please cite it:
Wu, C.-I. (2026). Academic Research Skills for Claude Code (Version 3.21.0) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.20696614