diff --git a/.github/CONTRIBUTING.md b/.github/CONTRIBUTING.md new file mode 100644 index 0000000..dee4c66 --- /dev/null +++ b/.github/CONTRIBUTING.md @@ -0,0 +1,46 @@ +# Contributing + +Thanks for helping. The bar: KERNEL stays a notebook you can open as one HTML file, with no server and no account, and the agent never does anything the person can't see, stop, or undo. + +## Layout + +- `src/` is the source of the two agent pages. `scripts/build.mjs` builds them into `docs/kernel-agent.html` and `docs/kernel-agent-mobile.html`, so never edit those two by hand. [`src/README.md`](../src/README.md) maps the tree: which folder holds the notebook, the agent, the Python harness, the CSS, and the shared markup. +- `docs/` is the GitHub Pages site. `kernel.html` (the notebook without the agent), `index.html`, the service worker, and the icons are edited directly. +- `skill/` is the `kernel-notebooks` Claude skill. `docs/kernel-notebooks.skill` is the same folder zipped for Claude.ai uploads (the site serves it), and CI checks that they match. +- `tests/verify_*.mjs` are static and fixture checks. `tests/e2e/` drives both builds in Chromium against a mock model provider. `tests/lib/source.mjs` has the helpers for reading the app's code. +- `examples/` holds curated agent runs, checked by `tests/verify_examples.mjs`. +- `specs/` holds the agent's contracts. `specs/agent-v2.4.md` is current, on top of the earlier ones. + +## Before you open a PR + +```bash +npm ci +npm run format # formats src/ with Prettier and rebuilds docs/ +npm run check # docs/ matches src/ and src/ is formatted (CI runs this) +npm test # static and fixture checks +npx playwright install chromium # once +npm run test:e2e # both builds in a real browser, a few minutes +``` + +Then look at your change in a browser: `python3 -m http.server -d docs` and open `http://localhost:8000/kernel-agent.html`. Check a phone width too (the mobile build is `kernel-agent-mobile.html`). For anything visible, put a before and after screenshot in the PR. If the change shows up in the demo, re-record it with `npm run demo`: it plays a scripted agent run against a local mock model and rewrites `docs/media/agent-demo.gif`, with no API key needed. + +## Writing tests + +- Test behavior, not text. Pull the functions you need out of a built page with `declarations(page, ...names)` or `functionSource(page, name)`, give them small stubs, and call them. `tests/verify_agent_v24.mjs` has many examples. +- When all you can check is that some code exists, use `has(page, snippet)`. It matches JavaScript tokens, so formatting never breaks it. +- UI behavior goes in `tests/e2e/ui.mjs`, which runs on both builds. Agent behavior goes in `tests/e2e/agent-loop.mjs`, against the mock provider in `tests/e2e/mock-provider.mjs`. + +## Rules of thumb + +- **One file, no server.** A built page loads only from the CDNs it already uses (Pyodide, KaTeX, Mermaid, fonts). Adding a runtime dependency needs a strong reason. +- **Keys stay in the browser** and go only to the API base the person chose. Exports and diagnostics never include them, and tests check this. +- **Desktop and mobile share code.** Shared behavior goes in `src/agent/js/` or `src/agent/markup/`; only phone-specific layout goes in `mobile.css` or `js/mobile/`. +- **Keep the agent's contracts.** Durable runs, checkpoints, the completion check, and context exclusions are specified in `specs/`. If a change alters one, update the current spec in the same PR. +- **Plain language in the UI.** Say what happened and what to do next. Offer Undo rather than a confirmation dialog when the action can be undone. + +## Releasing + +1. Bump `version` in `package.json` and run `npm run build`. The titles, the brand, and export metadata pick it up. +2. If the mobile app shell changed, bump the cache name in `docs/kernel-agent-sw.js` and the test that checks it. +3. Add a section to `CHANGELOG.md`. +4. Tag `vX.Y.Z` and push the tag, or run the Release workflow. It checks that the tag matches `package.json`, runs the checks, and publishes a GitHub Release. The release uses the changelog section as notes and attaches the single-file pages and the skill. diff --git a/.github/ISSUE_TEMPLATE/bug.yml b/.github/ISSUE_TEMPLATE/bug.yml new file mode 100644 index 0000000..276a6ac --- /dev/null +++ b/.github/ISSUE_TEMPLATE/bug.yml @@ -0,0 +1,31 @@ +name: Bug +description: Something broke or behaved wrongly in KERNEL, the agent, or the phone build +labels: [bug] +body: + - type: dropdown + id: app + attributes: + label: Which app + options: + - KERNEL (notebook) + - KERNEL·A (agent, desktop) + - KERNEL·M (agent, phone) + validations: { required: true } + - type: textarea + id: what + attributes: + label: What happened, and what you expected + description: Steps to reproduce if you have them. A screenshot helps for anything visual. + validations: { required: true } + - type: textarea + id: agent + attributes: + label: Agent details (if the agent was involved) + description: >- + Provider and model. More (⋯) → Export diagnostics saves a redacted file with versions and run + events, never prompts, code, file contents, or keys; attach it if you can. + - type: input + id: browser + attributes: + label: Browser and OS + placeholder: "Chrome 140 on macOS, Safari on iOS 19, …" diff --git a/.github/ISSUE_TEMPLATE/config.yml b/.github/ISSUE_TEMPLATE/config.yml new file mode 100644 index 0000000..0086358 --- /dev/null +++ b/.github/ISSUE_TEMPLATE/config.yml @@ -0,0 +1 @@ +blank_issues_enabled: true diff --git a/.github/ISSUE_TEMPLATE/idea.yml b/.github/ISSUE_TEMPLATE/idea.yml new file mode 100644 index 0000000..506b698 --- /dev/null +++ b/.github/ISSUE_TEMPLATE/idea.yml @@ -0,0 +1,14 @@ +name: Idea +description: Something KERNEL should do, or do better +labels: [idea] +body: + - type: textarea + id: want + attributes: + label: What you were trying to do + description: The analysis or task, and where KERNEL got in the way. + validations: { required: true } + - type: textarea + id: how + attributes: + label: What would help diff --git a/.github/release-footer.md b/.github/release-footer.md new file mode 100644 index 0000000..4ec0974 --- /dev/null +++ b/.github/release-footer.md @@ -0,0 +1,6 @@ + +### Use it + +- **In your browser:** [KERNEL](https://dogum.github.io/kernel/kernel.html) · [KERNEL·A, with the agent](https://dogum.github.io/kernel/kernel-agent.html) · [KERNEL·M, for phones](https://dogum.github.io/kernel/kernel-agent-mobile.html) +- **Offline or self-hosted:** download the `.html` files below. Each is the whole app; open it in a browser or put it on any static host. +- **The Claude skill:** upload `kernel-notebooks.skill` in Claude.ai under Settings → Capabilities → Skills. diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml new file mode 100644 index 0000000..9dbc763 --- /dev/null +++ b/.github/workflows/release.yml @@ -0,0 +1,64 @@ +name: Release + +# Two ways to release: +# 1. Push a tag vX.Y.Z (it must match "version" in package.json). +# 2. Run this workflow manually; it reads the version from package.json and tags the commit. +on: + push: + tags: ['v*'] + workflow_dispatch: + +permissions: + contents: write + +jobs: + release: + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v4 + - uses: actions/setup-node@v4 + with: + node-version: 22 + cache: npm + + - name: Resolve version and tag + id: v + run: | + v=$(node -p 'require("./package.json").version') + if [ "${GITHUB_REF_TYPE}" = "tag" ]; then + [ "v$v" = "${GITHUB_REF_NAME}" ] || { echo "tag ${GITHUB_REF_NAME} != package.json v$v"; exit 1; } + fi + echo "version=$v" >> "$GITHUB_OUTPUT" + echo "tag=v$v" >> "$GITHUB_OUTPUT" + + - run: npm ci + - run: npm run check + - run: npm test + + - name: Package the skill and the single-file apps + run: | + set -euo pipefail + mkdir -p dist/kernel-notebooks + cp -r skill/. dist/kernel-notebooks/ + (cd dist && zip -qr kernel-notebooks.skill kernel-notebooks) + cp docs/kernel.html docs/kernel-agent.html docs/kernel-agent-mobile.html dist/ + + - name: Release notes from CHANGELOG + run: | + awk -v v="${{ steps.v.outputs.version }}" '/^## /{p = index($0, "[" v "]") > 0; next} p' CHANGELOG.md > notes.md + [ -s notes.md ] || echo "See CHANGELOG.md." > notes.md + printf '\n---\n' >> notes.md + cat .github/release-footer.md >> notes.md + cat notes.md + + - uses: softprops/action-gh-release@v2 + with: + tag_name: ${{ steps.v.outputs.tag }} + target_commitish: ${{ github.sha }} + name: ${{ steps.v.outputs.tag }} + body_path: notes.md + files: | + dist/kernel-notebooks.skill + dist/kernel.html + dist/kernel-agent.html + dist/kernel-agent-mobile.html diff --git a/.github/workflows/verify.yml b/.github/workflows/verify.yml index 9c90c97..e44bcb6 100644 --- a/.github/workflows/verify.yml +++ b/.github/workflows/verify.yml @@ -9,18 +9,21 @@ permissions: jobs: static: - name: Static and fixture checks + name: Build, format, and fixture checks runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: actions/setup-node@v4 with: node-version: 22 - - run: node scripts/sync_agent_builds.mjs --check - - run: node tests/verify_agent_v2.mjs - - run: node tests/verify_agent_v23.mjs - - run: node tests/verify_agent_v24.mjs - - run: node tests/verify_examples.mjs + cache: npm + - run: npm ci + - run: npm run check + - run: npm test + - name: docs/kernel-notebooks.skill matches skill/ + run: | + unzip -q docs/kernel-notebooks.skill -d "$RUNNER_TEMP/skill" + diff -r "$RUNNER_TEMP/skill/kernel-notebooks" skill browser: name: Browser end-to-end (desktop + mobile builds) @@ -31,6 +34,7 @@ jobs: - uses: actions/setup-node@v4 with: node-version: 22 - - run: npm install --no-save --no-package-lock playwright@1.56.1 + cache: npm + - run: npm ci - run: npx playwright install --with-deps chromium - - run: node tests/e2e/run.mjs + - run: npm run test:e2e diff --git a/CHANGELOG.md b/CHANGELOG.md new file mode 100644 index 0000000..7130519 --- /dev/null +++ b/CHANGELOG.md @@ -0,0 +1,58 @@ +# Changelog + +Notable changes to KERNEL, KERNEL·A (the agent), and KERNEL·M (the phone build). Versions follow [semantic versioning](https://semver.org/); the current one is `version` in `package.json`. + +## [Unreleased] + +### Getting started +- Example links: `kernel-agent.html?example=` opens one of the curated runs in `examples/` with its original outputs, in a new notebook if the current one has work in it. The welcome card has a *See a real agent run* button, and the landing page and README link all four. Fleet DNA reads public data the repository doesn't include, so it links to where to get it. +- A demo of an agent run at the top of the README and the landing page, recorded by `npm run demo`. +- The model chip in the agent composer stays on one line. + +### Fixed +- While a notebook switch is saving and restoring, cells, agent runs and restarts are refused with a message to wait, and example links wait for it, so nothing runs in, or is imported into, a half-switched notebook. +- An OpenAI or xAI tool call with malformed JSON arguments is refused instead of running with empty input. Before, a malformed `run_all` reran the whole notebook. (Reported by Codex review on #5.) + +### For contributors +- The agent pages are built from `src/` by `scripts/build.mjs`. The code is split into about 60 files: notebook, agent, startup, and mobile JavaScript, CSS, shared markup, and the Python harness as a real `.py` file. The build still produces one self-contained HTML file per page. +- The source is formatted with Prettier. Checking every file's syntax tree before and after showed the formatting changed no code. +- Tests read the app through a JavaScript parser (`tests/lib/source.mjs`) instead of regexes, so formatting can't break them. +- `npm run build | check | format | test | test:e2e`, plus a CONTRIBUTING guide, issue templates, and a release workflow. +- The version lives in `package.json` and is filled in at build time. +- A tidier root: the agent specs moved to `specs/` (with an index), the packaged skill to `docs/` (the site serves it for download), the contributing guide to `.github/`, and the Prettier settings into `package.json`. The README is shorter. + +## [2.4.0] - 2026-09-28 + +### Agent +- Prompts are stable enough for the provider to cache them: the system prompt and tools never change within a thread, and live notebook state goes in an append-only block on the newest message. +- Budgets count effective tokens: cache reads cost about a tenth. +- `add_cells` and `edit_cell` take `run: true`, so writing a cell and seeing its output is one tool call. +- Rate limits, overloads, dropped or stalled streams, and truncated responses retry automatically. +- Claude Opus 5.5 is the default, with adaptive thinking, effort control, and server-side fallbacks. + +### Notebook +- Python runs in a Web Worker, so a runaway cell never freezes the page. Interrupt it from the toolbar or with `i i`. +- Files are stored once by content hash, so checkpoints and ZIP exports no longer duplicate them. +- Undo instead of dialogs: deleting a cell, a variable, or a file, and clearing outputs, show an Undo toast. `z` restores the last deleted cell. +- A first-run card with a sample dataset. Files dropped anywhere on the page are mounted. One-click head, describe, missing values, correlations, value counts, and histograms from the variable inspector. Tables copy as TSV or download as CSV, and figures download as PNG. +- Errors offer *Fix with agent*, and tracebacks name the cell (`Cell 4, line 1`). +- The KERNEL wordmark opens the app menu (other KERNEL apps, project page, theme), replacing the floating dock. +- On phones, sheets clear the tab bar, SEND stays visible, and the cell toolbar no longer covers code. + +### Docs +- Four curated real-world agent runs in `examples/`. + +## [2.3.1] - 2026-09-03 +- Completion-safe autonomy: a run that used tools must finish its visible plan and pass the `finish_run` evidence check, so a model that simply stops calling tools can't claim success. + +## [2.3.0] - 2026-09-03 +- Durable runs with checkpoints, pause and resume, and recovery after reload. Cell and file lineage with stale-output detection. An artifact workspace. Exact checkpoints and forks. Portable `.kernel.zip` handoff. See [`specs/agent-v2.3.md`](specs/agent-v2.3.md). + +## [2.0.0] - 2026-09-03 +- KERNEL Agent v2: Anthropic, OpenAI, and xAI adapters, several threads per notebook, and portable workspaces. See [`specs/agent-v2.md`](specs/agent-v2.md). + +## 2026-06-19 +- KERNEL·M, the agent as an installable, offline-capable phone app. + +## [0.1.0] - 2026-06-07 +- KERNEL, a Python notebook in one HTML file, with the `kernel-notebooks` Claude skill and the GitHub Pages site. diff --git a/README.md b/README.md index 53c0059..44a52b5 100644 --- a/README.md +++ b/README.md @@ -1,154 +1,86 @@ -# KERNEL - -A complete Python notebook that runs entirely in your browser, from **one self-contained HTML file**. Pyodide under the hood — no install, no build step, no server, and nothing leaves your machine. - -**[→ Launch it / see it live](https://dogum.github.io/kernel/)** - -This repo bundles a few things that belong together: +
-1. **KERNEL** — the notebook itself (`docs/kernel.html`): a single HTML file you can open, host, or fork. -2. **`kernel-notebooks`** — a Claude skill for authoring exceptional notebooks *for this runtime*. -3. **KERNEL Agent v2.4** — a durable, multi-provider notebook agent that can plan, execute, recover, compare models, and carry a complete workspace between devices (`docs/kernel-agent.html`; architecture in [`AGENT-V24-SPEC.md`](AGENT-V24-SPEC.md) on the [`AGENT-V23-SPEC.md`](AGENT-V23-SPEC.md) foundation). -4. **KERNEL·M** — a mobile / PWA build of the Agent (`docs/kernel-agent-mobile.html`): touch-friendly, installable to the home screen, and offline-capable. -5. **Real agent runs** — four curated, reproducible examples spanning autonomous software construction, synthetic operations, uncertain systems modeling, and wide public-data analysis ([`examples/`](examples/)). +# KERNEL -## What KERNEL is +**A Python notebook in one HTML file, with an agent that writes, runs and checks the analysis with you.** -Open `kernel.html` and you have a working kernel: write Python and markdown cells, run them in execution order, and get back stdout, plots, interactive Plotly, rendered DataFrames, KaTeX math, and Mermaid diagrams. It includes a data workspace (drop in CSVs and files, preview them, insert a read snippet), a live variable inspector with click-to-expand detail, a composable panel layout that collapses back to a calm centered notebook, `.ipynb` round-trip, `.py` export, and local persistence. +No install, no server, no account. Python runs in your browser through Pyodide, and your data stays there. -It's client-only by design. The Python executes in your tab via Pyodide/WebAssembly; your data stays in the page. The only network it needs is the one-time Pyodide download (CDN, ~10 MB, cached after) and the CDN fonts. +KERNEL·A loads a sample sales dataset; asked which region and channel bring in the most revenue, the agent writes and runs a pivot table and a stacked bar chart, then summarizes what stands out. -## Install +[![verify](https://github.com/dogum/kernel/actions/workflows/verify.yml/badge.svg)](https://github.com/dogum/kernel/actions/workflows/verify.yml) +[![License](https://img.shields.io/badge/license-Apache--2.0-blue)](LICENSE) -### The notebook +**[Open KERNEL·A](https://dogum.github.io/kernel/kernel-agent.html)** · **[KERNEL without the agent](https://dogum.github.io/kernel/kernel.html)** · **[On your phone](https://dogum.github.io/kernel/kernel-agent-mobile.html)** · **[Real agent runs](#real-agent-runs)** · **[Changelog](CHANGELOG.md)** -There's nothing to install — it's one file. +
-- **Use it now:** open the [live page](https://dogum.github.io/kernel/) and click *Launch KERNEL*. -- **Run it locally:** download [`docs/kernel.html`](docs/kernel.html) and open it in a browser. -- **Host it yourself:** drop the file on any static host (it's already served from `/docs` via GitHub Pages here). +--- -### The skill +| | What it is | +|---|---| +| **KERNEL** · [`docs/kernel.html`](docs/kernel.html) | The notebook: Python and markdown cells, plots, interactive Plotly, DataFrames, KaTeX, Mermaid, a data workspace, a variable inspector, and `.ipynb` round-trip. | +| **KERNEL·A** · [`docs/kernel-agent.html`](docs/kernel-agent.html) | The notebook with a bring-your-own-key agent (Anthropic, OpenAI, xAI) that plans, writes and runs cells, sees text and figures, and recovers safely after interruption. | +| **KERNEL·M** · [`docs/kernel-agent-mobile.html`](docs/kernel-agent-mobile.html) | KERNEL·A for phones: bottom sheets and a tab bar, installable, and offline after the first load. | +| **`kernel-notebooks`** · [`skill/`](skill/) | A Claude skill for writing notebooks that make the most of this runtime. | -**Claude.ai** — download [`kernel-notebooks.skill`](kernel-notebooks.skill) (or grab it from the latest Release), then Settings → Capabilities → Skills → upload. +## Real agent runs -**Claude Code** — copy the `skill/` folder into your skills directory: +Four unedited runs, with their prompts, decisions, corrections and rough edges kept. **Open in KERNEL·A** loads the notebook with the outputs the run produced; nothing runs until you ask. -```bash -git clone https://github.com/dogum/kernel.git -cp -r kernel/skill ~/.claude/skills/kernel-notebooks -``` +| Run | What the agent did | | +|---|---|---| +| [Fleet electrification](examples/fleet-dna/) | Real, wide public data on commercial-vehicle duty cycles: data-quality forensics, leakage-safe modeling, three duty-cycle archetypes | [Open in KERNEL·A](https://dogum.github.io/kernel/kernel-agent.html?example=fleet-dna) | +| [Regex engine](examples/regex-engine/) | A from-scratch NFA/DFA engine, three repaired semantic bugs, 20,000 consecutive agreements with Python `re.fullmatch` | [Open in KERNEL·A](https://dogum.github.io/kernel/kernel-agent.html?example=regex-engine) | +| [Lunar settlement launches](examples/lunar-settlement/) | A bottom-up launch model with explicit assumptions: baseline 121 launches, P10/P50/P90 of 112/134/161 | [Open in KERNEL·A](https://dogum.github.io/kernel/kernel-agent.html?example=lunar-settlement) | +| [Ares Station operations](examples/ares-station/) | A 4,320-hour colony twin, seven diagnosed incidents, a maintenance model and a stress-tested operating policy | [Open in KERNEL·A](https://dogum.github.io/kernel/kernel-agent.html?example=ares-station) | -The folder containing `SKILL.md` is what Claude Code loads. +They were captured with KERNEL Agent 2.3.0. The published notebooks keep the visible cells and outputs, and drop the private chat and provider metadata, run identifiers and checkpoint duplicates. [`examples/`](examples/) has each run's prompt, usage and limitations. -**Anthropic API** — skills can be deployed org-wide via the API; see the [Claude docs](https://docs.claude.com). +## Use it -## How it's structured +**The apps** are single HTML files. Open them from the [live page](https://dogum.github.io/kernel/), download one from [`docs/`](docs/) and open it locally, or put it on any static host. Pyodide downloads about 10 MB on first run and is cached after. Installing KERNEL·M to a home screen and using it offline need https, which GitHub Pages provides. -``` -kernel.html ............... lives in docs/ (served live on GitHub Pages) -kernel-notebooks.skill .... packaged skill, ready to upload to Claude.ai -AGENT-SPEC.md ............. original Anthropic-only agent specification -AGENT-V2-SPEC.md .......... provider, thread, context and workspace architecture -AGENT-V23-SPEC.md ......... durable runs, lineage, artifacts and handoff architecture -AGENT-V24-SPEC.md ......... cache-stable prompts, interruptible worker, retries, blob storage -skill/ -├── SKILL.md .............. runtime contract, the live-in-the-loop + multimodal sections, -│ narrative craft, output discipline -├── references/ -│ ├── runtime.md ........ the hard runtime facts (display helpers, ordered output, -│ │ %pip vs import, the markdown feature matrix incl. KaTeX/Mermaid) -│ └── chartsmanship.md .. matplotlib house style, static vs interactive -└── scripts/ - └── build_notebook.py . assembles a valid .ipynb from a JSON cell spec -docs/ -├── index.html ................. the landing / launch page -├── kernel.html ................ the notebook -├── kernel-agent.html .......... the agentic notebook (bring your own key) -├── kernel-agent-mobile.html ... the mobile / PWA build of the agent -├── kernel-agent-sw.js ......... service worker (offline cache for the PWA) -└── .nojekyll -scripts/ -└── sync_agent_builds.mjs ..... syncs the shared desktop/mobile runtime -tests/ -├── verify_agent_v2.mjs ....... provider compatibility and v2 regression checks -├── verify_agent_v23.mjs ...... durability, lineage, safety and handoff checks -├── verify_agent_v24.mjs ...... caching, Claude adapter, retries, worker and storage checks -├── verify_examples.mjs ....... curated notebook shape, privacy and artifact checks -└── e2e/ ...................... Playwright suites that drive both builds against a mock provider -examples/ -├── ares-station/ ............. 180-sol colony operations intelligence system -├── regex-engine/ ............. from-scratch engine with differential testing -├── lunar-settlement/ ......... mass, transport and Monte Carlo launch model -└── fleet-dna/ ................ real public-data duty-cycle analysis -``` +**The agent** needs an API key from Anthropic, OpenAI, or xAI; add it in the agent's settings. Keys stay in your browser and requests go straight to the provider (or a compatible gateway you choose). -`SKILL.md` is the entry point and is always in context when the skill triggers; the references are pulled in only when relevant. - -## KERNEL Agent v2.4 - -KERNEL Agent turns the notebook into an exploratory-analysis workbench: you describe what you want, and an agent writes the markdown and code cells, runs them, **sees** the results (text *and* figures), and iterates with you in the loop. It is a client-only, bring-your-own-key design with first-class adapters for Anthropic, OpenAI, and xAI/Grok. Keys remain in browser storage and requests go directly to the API base you select. - -Launch it from the [live page](https://dogum.github.io/kernel/) or open [`docs/kernel-agent.html`](docs/kernel-agent.html). [`AGENT-V24-SPEC.md`](AGENT-V24-SPEC.md) is the current release contract, layered on [`AGENT-V23-SPEC.md`](AGENT-V23-SPEC.md). [`AGENT-V2-SPEC.md`](AGENT-V2-SPEC.md) records the provider/thread foundation, and [`AGENT-SPEC.md`](AGENT-SPEC.md) remains the historical v1 design and tool-contract background. - -What it does today, beyond the core loop: - -- **Cache-stable, cost-aware runs** — the system prompt and tools stay byte-identical for a thread; live notebook state rides on the newest message as an append-only `` block, and older rich results compress in fixed epochs, so long runs keep hitting provider prompt caches. Budgets count effective tokens (cache reads at 10%). -- **Add-and-run tools** — `add_cells` and `edit_cell` accept `run: true`, so writing or repairing a cell and seeing its outputs takes one tool call instead of two. -- **Resilient provider calls** — rate limits, overloads, server errors, dropped streams, and stalls retry automatically with backoff that honors `retry-after`; a response cut off by the output limit retries with a larger cap. -- **Claude Opus 5.5 by default** — adaptive thinking with explicit effort, summarized reasoning and progress updates in the transcript, thinking blocks replayed unchanged across tool steps, conversation-bound thinking that degrades safely, and server-side refusal fallbacks. -- **Interruptible Python** — Pyodide runs in a Web Worker, so a slow or runaway cell never freezes the page. Interrupt from the toolbar or with `i i`; agent-run cells have a time limit. On cross-origin-isolated hosts interrupts keep your variables; elsewhere KERNEL restarts Python and remounts your files. -- **Provider parity** — Anthropic uses the native Messages API; OpenAI and xAI use the Responses API with the same KERNEL tool loop, multimodal results, stop behavior, and token accounting. Responses calls set `store: false`, and encrypted reasoning continuation items are preserved locally when returned. -- **Live Markdown transcript** — narration streams token-by-token and renders headings, tables, code, math, and Mermaid when complete. Reasoning summaries use a separate progressive-disclosure panel; tool calls remain compact, click-to-cell action chips. -- **Multiple threads per notebook** — create, rename, switch, or delete independent threads without mixing notebooks. Full messages and transcripts live in IndexedDB rather than a size-capped localStorage string. -- **Durable runs and recovery** — every request has a persisted run, phase, event timeline, visible plan, token/tool/time budgets, safe pause/resume, and recovery after reload. An ambiguous interrupted mutation is never silently repeated. -- **Completion-safe autonomy** — a tool-using or planned run must complete its visible plan and pass the `finish_run` evidence contract. A provider merely stopping tool calls cannot create a false success; AUTO extends tool/time checkpoints only while durable progress continues, while the token budget remains a hard cost boundary. -- **Notebook intelligence** — KERNEL tracks conservative Python dependencies, cell/file/environment provenance, and `fresh`, `stale`, or `historical` output state. Structured tracebacks navigate back to the exact cell and source line. -- **Artifact workspace** — uploads, working files, and final results have stable IDs, safe folder paths, previews, lifecycle controls, provenance, and notebook isolation. Folder upload preserves relative paths and collisions never silently overwrite data. Payloads are stored once by content hash, so checkpoints and ZIP exports don't duplicate your files. -- **Exact checkpoints and forks** — runs checkpoint notebook cells, rendered outputs, artifacts, environment, and thread state. Restore in place or fork a new notebook from that exact point without changing the source. -- **Conversations that persist and travel** — one-click `.kernel.zip` export carries the notebook, outputs, all threads, artifacts, durable runs, usage, environment, and checkpoints; open it elsewhere to resume. A separate share-safe ZIP strips chat/run history and inputs, scans copied text, and includes approved final results only. -- **Compact private run handoff** — a checkpoint-free `.kernel-run.zip` preserves the current notebook, outputs, artifacts, complete threads, and run ledgers for debugging or review without repeatedly embedding every historical checkpoint. It remains private and unredacted. -- **Multimodal input** — paste or drag images into the chat (sketch a chart, screenshot a figure); the agent sees them. -- **Explicit context control** — input, output, cache, and reasoning tokens are exposed. Pin or exclude cells and artifacts; large results use bounded handles; compaction changes only the active API payload, never the full local thread. When shared runtime state cannot be isolated safely, agent execution/inspection stops rather than bypass an exclusion; human cell controls remain available. -- **Read-only model comparison** — send the same notebook-grounded prompt to up to six configured provider/model profiles and compare answers, latency, and reported token use without giving contenders mutation tools. -- **Notebook-isolated workspaces** — stable cell IDs, outputs, user uploads, and agent results are restored only with their notebook. Switching notebooks snapshots state and resets the Python namespace so data cannot leak across workspaces. -- **Faster output and variable surfaces** — inline/panel output changes move existing DOM nodes instead of rebuilding rich output; the variable panel adds deterministic name/type/memory sorting and explicit refresh. -- **Built for exploring** — a first-run card loads a sample dataset or your own files, and files dropped anywhere on the page are mounted. The variable inspector turns a DataFrame or Series into head, describe, missing-value, correlation, value-count, or histogram cells in one click. Tables copy as TSV or download as CSV, figures download as PNG, and long outputs collapse. Errors offer *Fix with agent*, and tracebacks name the notebook cell instead of an internal path. -- **Undo instead of dialogs** — deleting a cell, a variable, or a data file, or clearing outputs, shows a toast with Undo (`z` restores the last deleted cell). -- **Autonomy modes** — AUTO runs free with a Stop button; STEP gates execution behind Approve/Skip. Every cell has an *ai* button that drops a stable cell reference into the composer. -- **Redacted diagnostics** — export an allowlisted support bundle with versions, counters, and sanitized run events, never prompts, source, file contents, tool payloads, or provider credentials. - -Provider model discovery is available from settings. Custom API base URLs are supported for compatible gateways; if a provider blocks direct browser CORS, use a gateway you control and trust. - -Release checks: +**The skill:** in Claude.ai, download [`kernel-notebooks.skill`](https://dogum.github.io/kernel/kernel-notebooks.skill) and upload it under Settings → Capabilities → Skills. In Claude Code, copy the folder: ```bash -node scripts/sync_agent_builds.mjs --check -node tests/verify_agent_v2.mjs -node tests/verify_agent_v23.mjs -node tests/verify_agent_v24.mjs -node tests/verify_examples.mjs -# browser end-to-end (needs Playwright + Chromium) -npm install --no-save playwright && npx playwright install chromium && node tests/e2e/run.mjs +git clone https://github.com/dogum/kernel.git +cp -r kernel/skill ~/.claude/skills/kernel-notebooks ``` -GitHub Actions runs the same checks on every push and pull request. - -## Real-world examples +## What the agent does -[`examples/`](examples/) contains four captured KERNEL Agent 2.3.0 runs with their exact prompts, curated `.ipynb` results, representative figures, original final artifacts where recovered, provider-reported usage, and candid limitations. They include a 53-cell Mars-operations simulation, a self-repairing regex-engine build, a Monte Carlo lunar-settlement launch model, and a 61-cell Fleet DNA analysis over a 4,705 × 372 public dataset. +- **Works in the notebook like you would.** It writes markdown and code cells, runs them, reads text, tables and figures, and fixes what breaks. You can paste or drop images into the chat, and every cell has an *ai* button that hands it to the agent. +- **Runs you can trust.** Every run is durable: it has a visible plan, budgets, pause and resume, and recovery after a reload. A run that used tools finishes only when its plan is done and a `finish_run` check confirms the evidence. AUTO runs freely with a Stop button; STEP asks before each run. +- **Checkpoints and forks.** Restore a notebook to any checkpoint, or fork a new notebook from it without touching the original. +- **Cheap and resilient.** Prompts stay stable enough for provider caching, budgets count cache reads at their real cost, and rate limits, overloads and dropped streams retry automatically. +- **Python that never freezes the page.** Pyodide runs in a worker. Interrupt a cell from the toolbar or with `i i`, and agent-run cells have a time limit. +- **Knows what is stale.** KERNEL tracks dependencies between cells and files, marks outputs as fresh, stale or historical, and links tracebacks to the cell and line. +- **Your files and context, under your control.** Uploads and results get stable IDs, previews and provenance. You can pin or exclude cells and files from the agent's context and see every token it uses. +- **Portable.** A `.kernel.zip` carries the notebook, outputs, threads, files, runs and checkpoints to another browser. A share-safe export strips history and keeps only approved results. +- **Several threads per notebook, and model comparison.** Send one question to up to six provider and model setups and compare the answers side by side. +- **Built for exploring.** It includes a sample dataset and drop-anywhere file mounting. The variable inspector adds cells in one click: head, describe, missing values, correlations, value counts, histograms. You can copy or download tables and figures, and *Fix with agent* appears on errors. Destructive actions offer Undo. -The published notebooks retain visible cells and outputs while removing private chat/provider metadata, internal run identifiers, browser fingerprints, and checkpoint duplication. Full workspace ZIPs remain recovery artifacts rather than repository examples. +Claude Opus 5.5 is the default model, with adaptive thinking. The design and every contract are in [`specs/`](specs/). -## Mobile / PWA (KERNEL·M) +## Privacy -`docs/kernel-agent-mobile.html` is a phone-friendly build of the Agent. The desktop layout is rebuilt as a single column: the notebook fills the screen and the Agent, Files and Variables panels become bottom sheets driven by a bottom navigation bar (swipe a sheet down to dismiss). It's also a Progressive Web App — installable to the home screen with its own icon, running standalone, and (served over https) caching the app shell and the Pyodide runtime through a service worker so it keeps working offline after the first load. It shares everything else with the Agent, including bring-your-own-key. +The notebook is client-side: Python runs in your browser, and your code and data leave the page only when you use the agent. The agent sends the context you allow straight to the provider or gateway you chose, with a key stored in this browser; OpenAI and xAI requests ask the provider not to store responses. Exports never include keys or provider settings. A full workspace export is lossless, so it keeps your prompts, cells and files as they are; check a share-safe export's redaction report before passing it on. -Open it from the [live page](https://dogum.github.io/kernel/) or [`docs/kernel-agent-mobile.html`](docs/kernel-agent-mobile.html). Install and offline need https (GitHub Pages provides it); opening the raw file over `file://` gives the responsive layout but not the service worker. +## Repository -## Privacy +| Folder | Holds | +|---|---| +| [`docs/`](docs/) | The GitHub Pages site: the three apps, the landing page, the service worker, and the packaged skill. | +| [`src/`](src/README.md) | Source of the two agent pages, built into `docs/` by `scripts/build.mjs`. | +| [`specs/`](specs/) | The agent's design and contracts. | +| [`skill/`](skill/) | The `kernel-notebooks` Claude skill. | +| [`examples/`](examples/) | The curated agent runs above. | +| [`tests/`](tests/) | Static and fixture checks, and browser tests that drive both builds against a mock model. | -The notebook is fully client-side: Python runs in your browser, and your code and data never leave the page unless you use the agent. The agent sends the selected active context directly to the provider/API base you choose (Anthropic, OpenAI, xAI, or a compatible gateway) using a key stored in this browser. Responses requests explicitly disable provider-side response storage where the protocol supports it. Exports never serialize provider configuration or stored keys; because a full workspace is intentionally lossless, user-authored prompts/cells/files are preserved verbatim. Inspect a share-safe archive's redaction report before redistributing it. +[`.github/CONTRIBUTING.md`](.github/CONTRIBUTING.md) has the commands to run before a PR and how releases work. Changes are in [`CHANGELOG.md`](CHANGELOG.md). ## License diff --git a/docs/index.html b/docs/index.html index 492177d..a55460d 100644 --- a/docs/index.html +++ b/docs/index.html @@ -75,6 +75,9 @@ /* launch cards */ .cards{display:grid; grid-template-columns:1fr 1fr; gap:1rem} +.demo{margin:2.4rem 0 0} +.demo img{display:block; width:100%; height:auto; border:1px solid var(--line-2); border-radius:10px; box-shadow:0 24px 60px -34px rgba(20,22,30,.45); background:var(--surface)} +.demo figcaption{font:13px/1.5 var(--mono); color:var(--ink-3); margin-top:.8em} @media (max-width:680px){.cards{grid-template-columns:1fr}} .card{border:1px solid var(--line); border-radius:12px; background:var(--surface); padding:1.4rem 1.5rem; display:flex; flex-direction:column} .card h3{font:700 1.15rem var(--disp); margin:0 0 .2em; display:flex; align-items:center; gap:.6em} @@ -123,6 +126,10 @@

A Python notebook that lives in one HTML file.

View source

Pyodide downloads ~10 MB of WebAssembly on first run (cached after). Desktop browser recommended.

+
+ KERNEL·A loads a sample sales dataset; asked which region and channel bring in the most revenue, the agent writes and runs a pivot table and a stacked bar chart, then summarizes what stands out. +
The agent answering a question about the sample dataset. It writes the cells, runs them, reads the results, and says what they show.
+
@@ -138,7 +145,7 @@

KERNEL Live

KERNEL Agent v2.4 Live

A durable notebook agent that plans, writes and runs cells, sees results, tracks lineage, and recovers safely after interruption. Python runs in an interruptible worker, prompts stay cache-friendly across long runs, and transient provider errors retry on their own. Bring an Anthropic, OpenAI, or xAI key; threads, artifacts, runs, checkpoints, outputs, and environment travel as one portable workspace.

Launch the Agent → -

Bring your own key · Anthropic / OpenAI / xAI · read the spec

+

Bring your own key · Anthropic / OpenAI / xAI · read the spec

KERNEL·M Live

@@ -149,6 +156,34 @@

KERNEL·M Live

+
+

Real agent runs

open one in your browser
+

Four unedited runs, kept with their prompts, decisions, corrections and rough edges. Opening one loads the notebook with the outputs it produced; nothing runs until you ask.

+
+
+

Fleet electrification

+

Real, wide public data on commercial-vehicle duty cycles: dictionary interpretation, data-quality forensics, leakage-safe modeling, and three duty-cycle archetypes with engineering recommendations.

+ Open in KERNEL·A → +
+
+

Regex engine

+

Software built from scratch in the notebook: an NFA/DFA engine, three semantic bugs found and repaired, and 20,000 consecutive agreements with Python's re.fullmatch.

+ Open in KERNEL·A → +
+
+

Lunar settlement launches

+

Assumption management and Monte Carlo uncertainty: a bottom-up launch model with a baseline of 121 launches and P10/P50/P90 of 112/134/161.

+ Open in KERNEL·A → +
+
+

Ares Station operations

+

A long open-ended simulation: a 4,320-hour colony twin, seven diagnosed incidents, a time-aware maintenance model, and a stress-tested operating policy.

+ Open in KERNEL·A → +
+
+

Prompts, run notes and figures for each are in examples/.

+
+

What's in the box

capabilities
    @@ -169,7 +204,7 @@

    KERNEL·M Live

    The kernel-notebooks skill

    for Claude

    A Claude skill that teaches the model to author exceptional notebooks for this runtime — well-structured, narrated analyses that use exactly the capabilities KERNEL has and none it doesn't. It's also the system prompt behind the agentic build: how to run the loop, read text and image outputs back, and explore with a human in the loop.

      -
    1. Claude.ai: download kernel-notebooks.skill, then Settings → Capabilities → Skills → upload.
    2. +
    3. Claude.ai: download kernel-notebooks.skill, then Settings → Capabilities → Skills → upload.
    4. Claude Code: copy the skill/ folder into your skills directory:
    git clone https://github.com/dogum/kernel.git
    diff --git a/docs/kernel-agent-mobile.html b/docs/kernel-agent-mobile.html
    index 5dea2b2..77b0566 100644
    --- a/docs/kernel-agent-mobile.html
    +++ b/docs/kernel-agent-mobile.html
    @@ -1,4 +1,5 @@
     
    +
     
     
     
    @@ -466,7 +467,7 @@
     /* ── v7 polish · in-app prompt dialog ── */
     .prompt-input{width:100%;box-sizing:border-box;font-family:var(--mono);font-size:13px;color:var(--ink);background:var(--surface-2);border:1px solid var(--line-2);border-radius:3px;padding:9px 11px;outline:none;caret-color:var(--accent)}
     .prompt-input:focus{border-color:color-mix(in srgb,var(--accent) 45%,var(--line-2))}
    -.prompt-acts{display:flex;justify-content:flex-end;gap:8px;margin-top:14px}
    +.prompt-acts{display:flex;justify-content:flex-end;gap:8px;margin-top:14px;flex-wrap:wrap}
     
     /* ===== KERNEL·A agent panel ===== */
     .brand-a{color:var(--accent);font-weight:800;letter-spacing:0}
    @@ -525,7 +526,8 @@
     .ag-mode{font-family:var(--mono);font-size:9px;font-weight:600;letter-spacing:.12em;color:var(--ink-3);
       border:1px solid var(--line-2);border-radius:3px;background:var(--surface);padding:5px 8px;cursor:pointer}
     .ag-mode.step{color:var(--accent);border-color:color-mix(in srgb,var(--accent) 45%,var(--line-2))}
    -.ag-steps{font-family:var(--mono);font-size:9px;letter-spacing:.08em;color:var(--ink-4)}
    +.ag-steps{font-family:var(--mono);font-size:9px;letter-spacing:.08em;color:var(--ink-4);white-space:nowrap}
    +#agModelChip{flex:0 1 auto;min-width:0;overflow:hidden;text-overflow:ellipsis;white-space:nowrap}
     .ag-usage{font-family:var(--mono);font-size:9px;letter-spacing:.06em;color:var(--ink-3);flex:1;text-align:right;cursor:pointer;white-space:nowrap;overflow:hidden;text-overflow:ellipsis}
     .ag-usage:hover{color:var(--accent)}
     .ag-stop{font-family:var(--mono);font-size:9.5px;font-weight:700;letter-spacing:.10em;color:var(--err);
    @@ -961,11 +963,13 @@ 

    Command mode

    @@ -4572,53 +12997,156 @@

    Command mode

    + + + + + + + + +
    +

    + +
    +
    Booting…
    +
    + + + + + + + + + + + + + + +
    +
    + +
    + + + +
    +
    +
    + + +
    +
    + + +
    + +
    + 0 cells + + + + Pyodide · loading + ⌘ shortcuts +
    + +
    + + + + + +
    + +
    + +
    + +
    + + + + + + + + + + + + + + diff --git a/src/agent/js/agent/010-overview.js b/src/agent/js/agent/010-overview.js new file mode 100644 index 0000000..750b030 --- /dev/null +++ b/src/agent/js/agent/010-overview.js @@ -0,0 +1,274 @@ +/* ════ KERNEL·A — cache-stable, interruptible, completion-safe provider-neutral agent (spec: specs/agent-v2.4.md on the specs/agent-v2.3.md foundation) ════ */ +var DOCS = + '# Authoring great in-browser Python notebooks\n\nThe target is a rich in-browser Python notebook runtime built on Pyodide (CPython on\nWebAssembly): one persistent kernel, Jupyter-style cells, last-expression echo, rich HTML\noutput, inline matplotlib, `%pip install`, and a small set of `display_*` helpers for\ninteractive output. (KERNEL is the reference implementation; everything here applies to any\nruntime that honors the same contract — see `references/runtime.md`.) Your job is to produce\nnotebooks that feel like a thoughtful human analyst wrote them — a narrated argument that\nhappens to be runnable — using exactly the capabilities the runtime has and none it doesn\'t.\nYou may be **assembling a finished notebook** in one pass, or **driving a live kernel turn by\nturn** with a human watching (see "Working live, in the loop"); the craft below applies to\nboth, but the live loop has its own discipline.\n\nA great notebook is not a script with comments. It is a sequence of small, rerunnable steps,\neach introduced by prose that says what is about to happen and why, followed by code,\nfollowed by an output worth looking at, followed by a sentence interpreting it. The reader\nshould be able to skim the markdown alone and understand the whole story.\n\n## The workflow\n\n1. **Clarify the goal and the inputs.** What question does the notebook answer? Is the input\n a dataset the user supplied, existing code to convert, a mounted file (via the **+Data**\n button → read from the working directory), or a synthetic dataset you generate? If you\n were given data, inspect its real columns/shape before writing analysis against guessed\n ones. If you were given code, see "Working from existing code or data" below.\n2. **Outline the narrative as section headers** before writing any code. A typical arc:\n framing → setup/imports → load & inspect → clean/feature-build → analysis/model →\n visualize → conclusion. Pick the arc that fits; don\'t pad.\n3. **Write the cells.** Alternate markdown and code. Keep each code cell to one idea. End\n cells on the value worth showing (see "Output discipline").\n4. **Pressure-test the result before you narrate it.** This is the difference between a\n demo and a real analysis. Before writing confident prose about a finding, confirm the\n finding actually exists: that the model converged, the cascade spread, the correlation is\n real, the clusters separate, the fit tracks the data. If the result is degenerate — a\n simulation that fizzles, an R² near zero, one cluster swallowing everything, a flat curve\n — **change the setup until there is real signal**, then narrate the true result. Never\n write a confident story over noise. If a result genuinely is null, say so plainly and\n show why; a clear negative result is honest, a fake positive is not.\n5. **Assemble a valid `.ipynb`** with the bundled script — never hand-write notebook JSON.\n Write a cell spec (a JSON list) and run:\n ```bash\n python scripts/build_notebook.py cells.json output.ipynb "Notebook Title"\n ```\n The script emits nbformat 4.5 the runtime imports cleanly. See the script header for the\n spec format.\n6. **Sanity-check the runtime fit.** Re-read every import and chart against\n `references/runtime.md` — no `requests`, no ipywidgets, no LaTeX in markdown, correct\n `%pip` vs. plain import. A notebook that errors on cell 1 is worthless. When you can run\n code, execute the cells in order (with `display_*` stubbed) to catch breakage before\n delivery.\n\n## Working live, in the loop\n\nWhen you are driving a *running* kernel turn by turn rather than assembling a finished file,\nthe job is exploratory analysis with a human watching. The rhythm is a loop, not a one-shot:\n\n1. **Propose one coherent step** — a markdown framing cell and the code cell it sets up, not\n ten cells you haven\'t seen run. Small steps keep you and the human oriented.\n2. **Run it and read what comes back.** Outputs return *in execution order* as a mix of:\n stream text (stdout/stderr), **figures as PNG images you can actually see**, rich\n tables/HTML rendered to text, and tracebacks. Look before you leap.\n3. **Interpret, then decide the next step from the evidence** — not from what you assumed the\n data would say. Drop a one-sentence takeaway in a markdown cell, then propose the next move.\n4. **Loop with the human.** Surface forks ("we could model this two ways…"), pause before\n expensive or destructive steps, and let them redirect. You are a pair-analyst, not an\n autopilot. The pressure-test rule (step 4 of the workflow) still holds: confirm a finding\n is real before you narrate it.\n\n### Reading outputs so you can actually see\n\nYou only "see" what a cell emits, so emit deliberately — the output *is* your sensory input:\n\n- **To judge a distribution, shape, fit, or trend, draw it.** A matplotlib / `display_plotly`\n figure comes back to you as an image; that is how you inspect a result visually. If you\n need to see it, plot it — then interpret what the picture shows.\n- **To reason over data, summarize it as text:** `df.head()`, `df.describe()`, `df.dtypes`,\n `value_counts()`, `df.shape`. **Never end a cell on a 10,000-row frame** — the whole thing\n streams back into your context and tells you little; a 10-row head tells you more.\n- **Read tracebacks and repair the cell** before continuing. A run that errors and barrels\n onward is worse than one that stops and fixes itself.\n- **Build on persistent state.** The kernel keeps everything in scope across cells (the\n Variables panel shows what\'s live) — reuse a frame you built three steps ago instead of\n recomputing it, and inspect an unfamiliar object before you assume its shape.\n\n## Working from existing code or data\n\nOften the input is not a blank prompt but **existing workbench code** or **a dataset** that\nshould become a notebook. Do not dump it into one cell. Convert it into a narrated notebook:\n\n- **Read it for intent first.** Identify the steps the code performs (load, transform,\n model, plot) and the story the data tells. The notebook\'s structure should follow that\n intent, not the original line order.\n- **Split into rerunnable steps**, one idea per cell, in the section arc above. Hoist\n imports to a single early cell. Pull magic numbers into named constants.\n- **Add the narration the code lacks**: a framing cell, a sentence of setup before each\n step, and a takeaway after each result. This is most of the value you add.\n- **Upgrade the outputs to the runtime.** Replace `print(df)` with a rendered DataFrame;\n turn a static chart into `display_plotly` when interactivity helps; typeset equations with\n `$…$` / `$$…$$` (KaTeX) and make a pipeline or state machine legible with a ```mermaid\n diagram instead of a paragraph describing it.\n- **Preserve behavior.** Don\'t silently change logic or results while restructuring; if you\n fix a real bug, call it out in prose.\n\n## Structure that reads well\n\n- **Open with a title cell**: an `#` H1, one or two sentences of framing, and a short\n bulleted list of what the reader will see. No code yet.\n- **One imports cell, early.** Group all imports there so first-run package loading happens\n once and later cells are fast. Set display options here too (`pd.set_option`, a seeded\n `np.random.default_rng`, matplotlib style).\n- **Section by section**: each analytical step gets a `##`/`###` header, 1–3 sentences of\n setup, the code, and then — crucially — a sentence of *takeaway* after the output. The\n takeaway is what separates a notebook from a transcript.\n- **Close with a conclusion cell** that states what was found, the caveats, and one or two\n next steps. This is the part readers remember.\n\n### The rhythm, concretely\n\nThree consecutive cells showing the markdown → code → takeaway beat:\n\n> **(markdown)** `## 2. Does engagement predict retention?`
    \n> We regress 30-day retention on first-week engagement, controlling for cohort size.\n>\n> **(code)** ```...fit model...; coef_table```  (cell ends on the rich table, not `print`)\n>\n> **(markdown)** Engagement is the dominant signal (β = 0.41, p < 0.001); cohort size\n> barely matters once it\'s included — so onboarding, not scale, is the lever.\n\nEvery section repeats that beat. No orphan code cells, no walls of prose.\n\n## Output discipline (this is where ordering matters)\n\nThe runtime builds each cell\'s output as a single ordered stream: prints, figures,\n`display_*` calls, and the cell\'s final value appear in the exact order they were produced.\nUse that.\n\n- **Let the last expression render.** End a cell on `df.head()`, a `Series`, a metrics\n table, or a fitted-model summary — the runtime shows its rich `_repr_html_` (DataFrames\n become real tables). Do **not** wrap it in `print()`, which collapses it to plain text.\n- **`print()` for narration of values mid-cell** (shapes, counts, chosen parameters); the\n rich result for the headline object at the end.\n- **`display(a, b, c)`** when you need to show several objects in order within one cell.\n- **Figures auto-display.** Just build the figure; it\'s captured inline. `plt.show()` is\n optional and harmless — figures created before the final value appear before it.\n- A cell that both plots and returns a table renders **plot then table**, matching execution\n order — put the plot code before the final expression if that\'s the order you want.\n\n## Visualization\n\nDefault to **static matplotlib** for clarity and print-fidelity; reach for **interactive**\nwhen hovering, zooming, or panning genuinely helps the reader explore.\n\n- Every chart: a title, axis labels **with units**, a legend when there is more than one\n series, a sensible `figsize`, and `fig.tight_layout()`. A chart without labels is a\n failure regardless of how pretty it is.\n- One idea per chart. If you\'re tempted to put five series on one axis, ask whether two\n small-multiples would read better.\n- **Interactive Plotly** is a one-liner: `%pip install plotly` once, then\n `display_plotly(fig, height=520)` — a fully interactive chart (zoom, pan, hover, lasso) in\n a sandboxed frame.\n- **Custom interactive HTML/JS** (d3, a hand-built widget, a Bokeh/Altair export):\n `display_html(html_string, height=520)` drops trusted HTML into a sandboxed iframe.\n- See `references/chartsmanship.md` for concrete recipes, a matplotlib house style, and the\n Plotly / `display_html` patterns.\n\n## Code craftsmanship\n\n- Idiomatic, readable Python: clear names, small functions, vectorized pandas/numpy over\n Python loops. A notebook is read more than it is run.\n- Seed every random process (`np.random.default_rng(SEED)`) so the narrative is reproducible\n across reruns.\n- Comment the *why*, not the *what*. The prose cell already said what; the comment explains a\n non-obvious choice.\n- Prefer composing state across cells (the kernel is persistent) over recomputing — but keep\n each cell independently rerunnable given the cells above it.\n\n## Hard runtime facts (do not violate)\n\nThese come from how the Pyodide runtime actually behaves. Full detail in\n`references/runtime.md`.\n\n- **Bundled scientific stack — just `import`:** numpy, pandas, scipy, scikit-learn,\n matplotlib, sympy, networkx, statsmodels, Pillow, beautifulsoup4, lxml, regex. The runtime\n auto-loads these on first import (the heavy ones — scipy/sklearn/statsmodels — take a\n moment the first time).\n- **Pure-Python extras — `%pip install ` first:** e.g. plotly, altair, humanize.\n micropip pulls them from PyPI.\n- **Will NOT work:** `requests` and raw sockets/threads/`subprocess` (no real OS/network\n layer); ipywidgets and `%matplotlib widget` (no widget comm — use `display_html` /\n `display_plotly` instead); any C-extension package Pyodide hasn\'t built (torch, tensorflow,\n polars, etc.).\n- **Markdown renders:** headings, **flat** bullet/numbered lists, GFM tables, blockquotes,\n fenced code (Python is syntax-highlighted), inline\n `code`/**bold**/*italic*/~~strike~~/links/images, **LaTeX math via KaTeX** (`$…$` inline,\n `$$…$$` display), and **Mermaid diagrams** inside a ```mermaid fenced block. It does **NOT**\n nest lists. Use `$…$` / `$$…$$` freely for equations, and reach for ```mermaid for\n flowcharts, sequence/state/ER diagrams, and pipeline schematics. (Reserve matplotlib\n mathtext for labels *inside* a chart, where KaTeX can\'t reach — use `\\frac`, not `\\dfrac`,\n there.)\n- **Data files** mounted via **+Data** land in the working directory; read them with\n `open("name")` / `pd.read_csv("name")`.\n\n## Anti-patterns (rewrite if you catch these)\n\n- One giant cell doing everything → split into rerunnable steps.\n- `print(df)` for a DataFrame → end the cell on the DataFrame so it renders as a table.\n- A chart with no title/labels → add them; state units.\n- Narrating a result you never checked → a fizzled sim or near-zero R² dressed up as a\n finding. Pressure-test first; fix the setup or report the null honestly.\n- `import requests` / network calls that assume a server → generate or mount the data.\n- `$\\sum_i x_i$` rendered as a matplotlib *figure* when the markdown cell would typeset it →\n use `$…$` / `$$…$$` directly; reserve mathtext for labels inside a chart.\n- Walls of markdown with no code, or walls of code with no narration → alternate.\n- Skipping the takeaway sentence after a result → always interpret what was just shown.\n\n## Reference files\n\n- `references/runtime.md` — the full runtime/Pyodide contract: display helpers (exact\n signatures), output ordering, the `%pip` vs import rule, package availability, the markdown\n feature matrix, and gotchas.\n- `references/chartsmanship.md` — a matplotlib house style, small-multiples and twin-axis\n recipes, the `display_plotly` and `display_html` patterns with complete examples, and how\n to choose static vs. interactive.\n- `scripts/build_notebook.py` — assembles a valid nbformat 4.5 `.ipynb` from a JSON cell\n spec. Always build notebooks with this rather than emitting JSON by hand.\n\n## Chart house style (condensed)\n- Every chart: title, axis labels WITH units, legend when >1 series, sensible figsize (~8x4.5), fig.tight_layout().\n- One idea per chart; prefer two small multiples over five series on one axis.\n- Seed RNGs; annotate the takeaway point rather than describing it only in prose.\n- Light grid only when it aids reading; no chartjunk; colorblind-safe default cycle.\n- Interactivity only when hover/zoom genuinely helps: %pip install plotly once, then display_plotly(fig, height=520).\n- Custom HTML/JS output goes through display_html(html, height) — sandboxed iframe.\n- matplotlib mathtext only for math inside a chart; markdown KaTeX ($…$/$$…$$) for narrative equations.\n\n# In-browser Python runtime contract\n\nEverything a notebook author needs to know about how the runtime executes code and renders\noutput. The runtime is CPython on WebAssembly (Pyodide) running entirely in the browser, with\none persistent kernel shared across all cells. (KERNEL is the reference implementation; the\nnames below — `display_html`, `display_plotly`, the ordering rules — are the contract any\ncompatible runtime should honor.)\n\n## Table of contents\n1. Execution model\n2. Output ordering\n3. Display helpers (exact signatures)\n4. matplotlib\n5. Packages: import vs `%pip`\n6. The markdown feature matrix\n7. Data files\n8. Things that don\'t exist here\n9. Performance notes\n\n## 1. Execution model\n\n- **One persistent namespace.** Variables, imports, and functions defined in one cell are\n available in every later cell, exactly like Jupyter. Build state up incrementally.\n- **Top-level `await` is allowed.** Cells run with `eval_code_async`, so\n `await something()` works at cell top level.\n- **Last-expression echo.** If a cell\'s final statement is an expression, its value is\n displayed: rich HTML via `_repr_html_` when the object provides it (pandas DataFrame,\n Series, styled objects), otherwise `repr()`. A trailing `;` or ending on an assignment\n suppresses the echo.\n- **`_`** holds the last echoed value, as in a REPL.\n- **Errors** show a trimmed traceback starting at the user\'s cell frame.\n\n## 2. Output ordering\n\nEach cell\'s outputs are emitted as a single **ordered** stream in execution order:\nstdout/stderr fragments, figures, `display_*` calls, and finally the cell\'s return value.\n\nPractical consequences:\n- `print("loading...")` then a plot then a returned table renders as: text, figure, table.\n- Figures created without an explicit `show()` are captured **right before** the cell\'s\n final value, so they appear above a returned table.\n- `plt.show()` / `fig.show()` are intercepted to flush pending text and capture the figure\n *at that point* in the stream — so multiple show() calls interleaved with prints render\n in the right order. They never emit the Agg "non-interactive" warning.\n\n## 3. Display helpers (exact signatures)\n\nThese are injected into every cell\'s namespace — no import needed.\n\n```python\ndisplay(*objs)\n# Render one or more objects inline, in order. Uses _repr_html_ when available,\n# else repr(). Use when you want to show several things from one cell.\n\ndisplay_html(html: str, height: int = 480)\n# Render a trusted HTML string in a sandboxed