Harness KitsAI workspace setup guides

Field notes · September 2026

Worked Examples

Real runs behind the kits. For each one: what we asked, what the agent did, what we checked by hand, and what went wrong.

Runs
6
Model
Local
Checked
By hand

How we ran them

OpenCode 1.18 and Pi 0.87 on a Mac, connected to the NaviGator development endpoint with NaviGator's local models: meta-muse-glimmer-30b for research and gpt-oss-120b for code. Results on the production service should be similar but can differ.

What we changed

To run unattended, commands were approved automatically. In your sessions the agent asks first. We used each kit's config and house rules as published, and we say when we changed them.

What we checked

Every number, citation, and code change below was checked against the source: PubMed, Crossref, the original PDF, or our own script. Time and token counts are from the agent's own logs.

Analyzing sensitive survey data

OpenCodeglimmer-30bKit 01

A theme summary of 40 coded interview records (fake data: names, ages, counties, themes, quotes), in the locked-down workspace from Kit 01.

Prompt
Summarize data/interviews-coded.csv: how many participants mention each theme, overall and by county. Write a short report to outputs/theme-summary.md that follows the workspace rules.
What it did
Wrote scripts/analyze_themes.py, ran it, and wrote the report. 39 seconds.
We checked
Every count matched our own tally. Every county cell under 5 was suppressed. No names, IDs, or quotes appeared in the output, and data/ was untouched.
Guardrails
With the rules file removed, we asked it to fetch a web page, search the web, run curl, edit a data file, and read /etc/hosts. All five were blocked by the config.
Watch for
The fake participant names appeared in OpenCode's session database because the agent read the CSV. Delete sessions when you're done. The model also retried the blocked curl five times before moving on.
Files
theme-summary.md · analyze_themes.py

Searching PubMed through an MCP server

OpenCodePiglimmer-30bKit 12

Five recent papers on rural telehealth, found with the PubMed MCP server and saved as a table.

Prompt
Use the pubmed MCP tools to find 5 peer-reviewed papers from 2020 or later on telehealth use in rural areas of the United States. Save a table to outputs/pubmed-results.md with PMID, DOI, title, first author, journal, and year. Only include details the tools returned; do not fill anything in from memory.
What it did
Three PubMed searches, one fetch of five full records, and the table. 14 seconds. In Pi, a single-record lookup through pi-mcp-adapter took 6 seconds.
We checked
All five PMIDs, DOIs, titles, first authors, and years matched PubMed's own records. Nothing was invented.
Watch for
Relevance was mixed: one result was a position paper, and two were only partly about rural care. Accurate metadata doesn't mean the paper belongs in your review.
Files
pubmed-results.md

A table from a government PDF to CSV

OpenCodeglimmer-30bKit 10

Table 1 of the Census Bureau report Health Insurance Coverage in the United States: 2023 (P60-284, 28 pages), extracted locally with the liteparse skill.

Prompt
Use the liteparse skill to extract Table 1 from sources/p60-284.pdf into outputs/table1.csv. Use one row per coverage type with columns: coverage_type, level, number_2022, moe_2022, number_2023, moe_2023, all in thousands as printed. Remove footnote markers from the labels. Record the PDF page and printed page number in outputs/table1-notes.md, along with the table's units and footnotes. Do the extraction with a script in scripts/.
What it did
Parsed the PDF on the Mac, found the table, and wrote a script that produced the CSV. About 3 minutes.
We checked
All 48 values (12 rows × 4 columns) matched the printed table. Footnote markers were removed from labels such as "Employment-based²", and both page numbers were right (PDF page 8, printed page 2).
Watch for
The footnotes in the notes file were garbled: text from the next column was mixed in. The run also used about 5 million input tokens, because the agent re-read parser output many times.
Files
table1.csv

Reading a paper with a local converter

OpenCodeglimmer-30bKit 12

The 15-page paper "Attention Is All You Need" converted to Markdown with the MarkItDown MCP server, then questioned about its results.

What it did
Converted the PDF, listed its numbered headings two levels deep (19 in all; it skipped the third-level 3.2.1–3.2.3), and reported both BLEU scores from the results table, quoting the table row. 81 seconds.
We checked
The scores (28.4 and 41.8) and the listed headings match the paper.
Watch for
Some lines lost the spaces between words ("EncoderandDecoderStacks"). The session used about 625,000 input tokens, because the full text went into the context more than once.

A practice scoping review, start to bibliography

OpenCodeglimmer-30bKit 13

Telehealth for adult mental health care in the rural United States, 2020–2025: searches in PubMed and OpenAlex, deduplication, title-and-abstract screening, PRISMA counts, and a BibTeX file. Five prompts in one workspace.

Steps of the practice review with time, what happened, and what we found
StepTimeResultWhat we found
Search53 s80 records, both searches loggedAll 40 PubMed abstracts were empty (a parser bug), and the date limit used entry date instead of publication date.
Check3.7 minBoth bugs found and fixedPubMed count went from 55 to 58, matching our own query. Its fix for one bad year was a special case, not a general rule.
Screen (first try)69 s19 include, 6 exclude, 54 unsureIt wrote keyword rules instead of reading: app matched "approach". Our house rules were too vague. We fixed them.
Screen (again)2.4 min13 include, 59 exclude, 7 unsure8 of 13 includes were right. 18 of 20 sampled excludes were defensible; one gave a false reason.
Bibliography39 s13 BibTeX entriesEvery DOI resolved and every title matched Crossref. Three used the online-first year.
Watch for
Twice the agent said "done" when the output was wrong. The check prompt in Kit 13, step 4, caught the first problem; reading the screening file caught the second.
Files
search-log.md · screening.csv · prisma-counts.md · included.bib

Fixing a bug in survey-scoring code

OpenCodegpt-oss-120bKit 20

A small Python package that scores Likert items. Two tests failed, and skipped answers (None) crashed the mean. We used the Plan agent first, then Build.

Prompt · Plan agent
Two tests in tests/test_scales.py fail. Find the cause. Also, survey respondents sometimes skip items; those arrive as None and scale_mean should ignore them. Propose a plan: the code changes and the new tests. Don't edit anything yet.
What it did
Plan (25 s): found that reverse_score ignored the scale minimum, and proposed skipping None, dividing by answered items, and three new tests. Build (29 s): made the changes, and all 6 tests passed. It didn't commit, as its house rules say.
We checked
The diff is correct. We tested extra cases (0–4 and 1–7 scales, all items skipped); all were right.
Watch for
Reading the diff found a gap the agent missed: scale_mean can't pass a scale minimum to reverse_score, so a reverse-scored item on a 0–4 scale is still wrong. Passing tests aren't the same as correct code.
Files
coding-fix.diff

What the runs have in common

These are the habits that caught every problem above.