Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
18 commits
Select commit Hold shift + click to select a range
96f811a
Build the agent pages from src/ modules
claude Sep 28, 2026
c19256c
Format the source with Prettier and test through a parser
claude Sep 28, 2026
d835167
Add CONTRIBUTING, CHANGELOG, issue templates, and a release workflow
claude Sep 28, 2026
49c0e8a
Show it working: demo recording, example links, README and landing page
claude Sep 28, 2026
8151391
Refuse Responses tool calls with malformed arguments
claude Sep 28, 2026
1a8c7e0
Tidy the repository root
claude Sep 28, 2026
2199641
Never import an example into a notebook the agent is working in
claude Sep 28, 2026
d1be313
Keep examples out of notebooks with unfinished runs or a busy kernel
claude Sep 28, 2026
11a8ed7
Lock out runs while notebooks switch; recheck before importing
claude Sep 28, 2026
62a2234
Keep cells and restarts from starting during a notebook switch
claude Sep 28, 2026
f7b13b3
Wait for startup restore before example links; hold cells during prep
claude Sep 28, 2026
ca8c4c6
Wait for notebook switches before opening an example
claude Sep 28, 2026
ff2ad66
Reset Python when an example reuses a blank notebook
claude Sep 28, 2026
a251ab7
Say when an example needs outside data to rerun
claude Sep 28, 2026
9895432
Note the notebook-switch and example-link fixes in the changelog
claude Sep 28, 2026
eced742
Wait for the startup agent load before example links decide
claude Sep 28, 2026
e736c0b
Keep examples out of notebooks with agent history
claude Sep 28, 2026
a642d8b
Recheck a new notebook before an example fills it
claude Sep 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
46 changes: 46 additions & 0 deletions .github/CONTRIBUTING.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# Contributing

Thanks for helping. The bar: KERNEL stays a notebook you can open as one HTML file, with no server and no account, and the agent never does anything the person can't see, stop, or undo.

## Layout

- `src/` is the source of the two agent pages. `scripts/build.mjs` builds them into `docs/kernel-agent.html` and `docs/kernel-agent-mobile.html`, so never edit those two by hand. [`src/README.md`](../src/README.md) maps the tree: which folder holds the notebook, the agent, the Python harness, the CSS, and the shared markup.
- `docs/` is the GitHub Pages site. `kernel.html` (the notebook without the agent), `index.html`, the service worker, and the icons are edited directly.
- `skill/` is the `kernel-notebooks` Claude skill. `docs/kernel-notebooks.skill` is the same folder zipped for Claude.ai uploads (the site serves it), and CI checks that they match.
- `tests/verify_*.mjs` are static and fixture checks. `tests/e2e/` drives both builds in Chromium against a mock model provider. `tests/lib/source.mjs` has the helpers for reading the app's code.
- `examples/` holds curated agent runs, checked by `tests/verify_examples.mjs`.
- `specs/` holds the agent's contracts. `specs/agent-v2.4.md` is current, on top of the earlier ones.

## Before you open a PR

```bash
npm ci
npm run format # formats src/ with Prettier and rebuilds docs/
npm run check # docs/ matches src/ and src/ is formatted (CI runs this)
npm test # static and fixture checks
npx playwright install chromium # once
npm run test:e2e # both builds in a real browser, a few minutes
```

Then look at your change in a browser: `python3 -m http.server -d docs` and open `http://localhost:8000/kernel-agent.html`. Check a phone width too (the mobile build is `kernel-agent-mobile.html`). For anything visible, put a before and after screenshot in the PR. If the change shows up in the demo, re-record it with `npm run demo`: it plays a scripted agent run against a local mock model and rewrites `docs/media/agent-demo.gif`, with no API key needed.

## Writing tests

- Test behavior, not text. Pull the functions you need out of a built page with `declarations(page, ...names)` or `functionSource(page, name)`, give them small stubs, and call them. `tests/verify_agent_v24.mjs` has many examples.
- When all you can check is that some code exists, use `has(page, snippet)`. It matches JavaScript tokens, so formatting never breaks it.
- UI behavior goes in `tests/e2e/ui.mjs`, which runs on both builds. Agent behavior goes in `tests/e2e/agent-loop.mjs`, against the mock provider in `tests/e2e/mock-provider.mjs`.

## Rules of thumb

- **One file, no server.** A built page loads only from the CDNs it already uses (Pyodide, KaTeX, Mermaid, fonts). Adding a runtime dependency needs a strong reason.
- **Keys stay in the browser** and go only to the API base the person chose. Exports and diagnostics never include them, and tests check this.
- **Desktop and mobile share code.** Shared behavior goes in `src/agent/js/` or `src/agent/markup/`; only phone-specific layout goes in `mobile.css` or `js/mobile/`.
- **Keep the agent's contracts.** Durable runs, checkpoints, the completion check, and context exclusions are specified in `specs/`. If a change alters one, update the current spec in the same PR.
- **Plain language in the UI.** Say what happened and what to do next. Offer Undo rather than a confirmation dialog when the action can be undone.

## Releasing

1. Bump `version` in `package.json` and run `npm run build`. The titles, the brand, and export metadata pick it up.
2. If the mobile app shell changed, bump the cache name in `docs/kernel-agent-sw.js` and the test that checks it.
3. Add a section to `CHANGELOG.md`.
4. Tag `vX.Y.Z` and push the tag, or run the Release workflow. It checks that the tag matches `package.json`, runs the checks, and publishes a GitHub Release. The release uses the changelog section as notes and attaches the single-file pages and the skill.
31 changes: 31 additions & 0 deletions .github/ISSUE_TEMPLATE/bug.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
name: Bug
description: Something broke or behaved wrongly in KERNEL, the agent, or the phone build
labels: [bug]
body:
- type: dropdown
id: app
attributes:
label: Which app
options:
- KERNEL (notebook)
- KERNEL·A (agent, desktop)
- KERNEL·M (agent, phone)
validations: { required: true }
- type: textarea
id: what
attributes:
label: What happened, and what you expected
description: Steps to reproduce if you have them. A screenshot helps for anything visual.
validations: { required: true }
- type: textarea
id: agent
attributes:
label: Agent details (if the agent was involved)
description: >-
Provider and model. More (⋯) → Export diagnostics saves a redacted file with versions and run
events, never prompts, code, file contents, or keys; attach it if you can.
- type: input
id: browser
attributes:
label: Browser and OS
placeholder: "Chrome 140 on macOS, Safari on iOS 19, …"
1 change: 1 addition & 0 deletions .github/ISSUE_TEMPLATE/config.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
blank_issues_enabled: true
14 changes: 14 additions & 0 deletions .github/ISSUE_TEMPLATE/idea.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
name: Idea
description: Something KERNEL should do, or do better
labels: [idea]
body:
- type: textarea
id: want
attributes:
label: What you were trying to do
description: The analysis or task, and where KERNEL got in the way.
validations: { required: true }
- type: textarea
id: how
attributes:
label: What would help
6 changes: 6 additions & 0 deletions .github/release-footer.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@

### Use it

- **In your browser:** [KERNEL](https://dogum.github.io/kernel/kernel.html) · [KERNEL·A, with the agent](https://dogum.github.io/kernel/kernel-agent.html) · [KERNEL·M, for phones](https://dogum.github.io/kernel/kernel-agent-mobile.html)
- **Offline or self-hosted:** download the `.html` files below. Each is the whole app; open it in a browser or put it on any static host.
- **The Claude skill:** upload `kernel-notebooks.skill` in Claude.ai under Settings → Capabilities → Skills.
64 changes: 64 additions & 0 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
name: Release

# Two ways to release:
# 1. Push a tag vX.Y.Z (it must match "version" in package.json).
# 2. Run this workflow manually; it reads the version from package.json and tags the commit.
on:
push:
tags: ['v*']
workflow_dispatch:

permissions:
contents: write

jobs:
release:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: npm

- name: Resolve version and tag
id: v
run: |
v=$(node -p 'require("./package.json").version')
if [ "${GITHUB_REF_TYPE}" = "tag" ]; then
[ "v$v" = "${GITHUB_REF_NAME}" ] || { echo "tag ${GITHUB_REF_NAME} != package.json v$v"; exit 1; }
fi
echo "version=$v" >> "$GITHUB_OUTPUT"
echo "tag=v$v" >> "$GITHUB_OUTPUT"

- run: npm ci
- run: npm run check
- run: npm test

- name: Package the skill and the single-file apps
run: |
set -euo pipefail
mkdir -p dist/kernel-notebooks
cp -r skill/. dist/kernel-notebooks/
(cd dist && zip -qr kernel-notebooks.skill kernel-notebooks)
cp docs/kernel.html docs/kernel-agent.html docs/kernel-agent-mobile.html dist/

- name: Release notes from CHANGELOG
run: |
awk -v v="${{ steps.v.outputs.version }}" '/^## /{p = index($0, "[" v "]") > 0; next} p' CHANGELOG.md > notes.md
[ -s notes.md ] || echo "See CHANGELOG.md." > notes.md
printf '\n---\n' >> notes.md
cat .github/release-footer.md >> notes.md
cat notes.md

- uses: softprops/action-gh-release@v2
with:
tag_name: ${{ steps.v.outputs.tag }}
target_commitish: ${{ github.sha }}
name: ${{ steps.v.outputs.tag }}
body_path: notes.md
files: |
dist/kernel-notebooks.skill
dist/kernel.html
dist/kernel-agent.html
dist/kernel-agent-mobile.html
20 changes: 12 additions & 8 deletions .github/workflows/verify.yml
Original file line number Diff line number Diff line change
Expand Up @@ -9,18 +9,21 @@ permissions:

jobs:
static:
name: Static and fixture checks
name: Build, format, and fixture checks
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- run: node scripts/sync_agent_builds.mjs --check
- run: node tests/verify_agent_v2.mjs
- run: node tests/verify_agent_v23.mjs
- run: node tests/verify_agent_v24.mjs
- run: node tests/verify_examples.mjs
cache: npm
- run: npm ci
- run: npm run check
- run: npm test
- name: docs/kernel-notebooks.skill matches skill/
run: |
unzip -q docs/kernel-notebooks.skill -d "$RUNNER_TEMP/skill"
diff -r "$RUNNER_TEMP/skill/kernel-notebooks" skill

browser:
name: Browser end-to-end (desktop + mobile builds)
Expand All @@ -31,6 +34,7 @@ jobs:
- uses: actions/setup-node@v4
with:
node-version: 22
- run: npm install --no-save --no-package-lock playwright@1.56.1
cache: npm
- run: npm ci
- run: npx playwright install --with-deps chromium
- run: node tests/e2e/run.mjs
- run: npm run test:e2e
58 changes: 58 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
# Changelog

Notable changes to KERNEL, KERNEL·A (the agent), and KERNEL·M (the phone build). Versions follow [semantic versioning](https://semver.org/); the current one is `version` in `package.json`.

## [Unreleased]

### Getting started
- Example links: `kernel-agent.html?example=<name>` opens one of the curated runs in `examples/` with its original outputs, in a new notebook if the current one has work in it. The welcome card has a *See a real agent run* button, and the landing page and README link all four. Fleet DNA reads public data the repository doesn't include, so it links to where to get it.
- A demo of an agent run at the top of the README and the landing page, recorded by `npm run demo`.
- The model chip in the agent composer stays on one line.

### Fixed
- While a notebook switch is saving and restoring, cells, agent runs and restarts are refused with a message to wait, and example links wait for it, so nothing runs in, or is imported into, a half-switched notebook.
- An OpenAI or xAI tool call with malformed JSON arguments is refused instead of running with empty input. Before, a malformed `run_all` reran the whole notebook. (Reported by Codex review on #5.)

### For contributors
- The agent pages are built from `src/` by `scripts/build.mjs`. The code is split into about 60 files: notebook, agent, startup, and mobile JavaScript, CSS, shared markup, and the Python harness as a real `.py` file. The build still produces one self-contained HTML file per page.
- The source is formatted with Prettier. Checking every file's syntax tree before and after showed the formatting changed no code.
- Tests read the app through a JavaScript parser (`tests/lib/source.mjs`) instead of regexes, so formatting can't break them.
- `npm run build | check | format | test | test:e2e`, plus a CONTRIBUTING guide, issue templates, and a release workflow.
- The version lives in `package.json` and is filled in at build time.
- A tidier root: the agent specs moved to `specs/` (with an index), the packaged skill to `docs/` (the site serves it for download), the contributing guide to `.github/`, and the Prettier settings into `package.json`. The README is shorter.

## [2.4.0] - 2026-09-28

### Agent
- Prompts are stable enough for the provider to cache them: the system prompt and tools never change within a thread, and live notebook state goes in an append-only block on the newest message.
- Budgets count effective tokens: cache reads cost about a tenth.
- `add_cells` and `edit_cell` take `run: true`, so writing a cell and seeing its output is one tool call.
- Rate limits, overloads, dropped or stalled streams, and truncated responses retry automatically.
- Claude Opus 5.5 is the default, with adaptive thinking, effort control, and server-side fallbacks.

### Notebook
- Python runs in a Web Worker, so a runaway cell never freezes the page. Interrupt it from the toolbar or with `i i`.
- Files are stored once by content hash, so checkpoints and ZIP exports no longer duplicate them.
- Undo instead of dialogs: deleting a cell, a variable, or a file, and clearing outputs, show an Undo toast. `z` restores the last deleted cell.
- A first-run card with a sample dataset. Files dropped anywhere on the page are mounted. One-click head, describe, missing values, correlations, value counts, and histograms from the variable inspector. Tables copy as TSV or download as CSV, and figures download as PNG.
- Errors offer *Fix with agent*, and tracebacks name the cell (`Cell 4, line 1`).
- The KERNEL wordmark opens the app menu (other KERNEL apps, project page, theme), replacing the floating dock.
- On phones, sheets clear the tab bar, SEND stays visible, and the cell toolbar no longer covers code.

### Docs
- Four curated real-world agent runs in `examples/`.

## [2.3.1] - 2026-09-03
- Completion-safe autonomy: a run that used tools must finish its visible plan and pass the `finish_run` evidence check, so a model that simply stops calling tools can't claim success.

## [2.3.0] - 2026-09-03
- Durable runs with checkpoints, pause and resume, and recovery after reload. Cell and file lineage with stale-output detection. An artifact workspace. Exact checkpoints and forks. Portable `.kernel.zip` handoff. See [`specs/agent-v2.3.md`](specs/agent-v2.3.md).

## [2.0.0] - 2026-09-03
- KERNEL Agent v2: Anthropic, OpenAI, and xAI adapters, several threads per notebook, and portable workspaces. See [`specs/agent-v2.md`](specs/agent-v2.md).

## 2026-06-19
- KERNEL·M, the agent as an installable, offline-capable phone app.

## [0.1.0] - 2026-06-07
- KERNEL, a Python notebook in one HTML file, with the `kernel-notebooks` Claude skill and the GitHub Pages site.
Loading
Loading