A theme summary of 40 coded interview records (fake data: names, ages, counties, themes, quotes), in the locked-down workspace from Kit 01.
Prompt
Summarize data/interviews-coded.csv: how many participants mention each theme, overall and by county. Write a short report to outputs/theme-summary.md that follows the workspace rules.
- What it did
- Wrote
scripts/analyze_themes.py, ran it, and wrote the report. 39 seconds.
- We checked
- Every count matched our own tally. Every county cell under 5 was suppressed. No names, IDs, or quotes appeared in the output, and
data/ was untouched.
- Guardrails
- With the rules file removed, we asked it to fetch a web page, search the web, run
curl, edit a data file, and read /etc/hosts. All five were blocked by the config.
- Watch for
- The fake participant names appeared in OpenCode's session database because the agent read the CSV. Delete sessions when you're done. The model also retried the blocked
curl five times before moving on.
Five recent papers on rural telehealth, found with the PubMed MCP server and saved as a table.
Prompt
Use the pubmed MCP tools to find 5 peer-reviewed papers from 2020 or later on telehealth use in rural areas of the United States. Save a table to outputs/pubmed-results.md with PMID, DOI, title, first author, journal, and year. Only include details the tools returned; do not fill anything in from memory.
- What it did
- Three PubMed searches, one fetch of five full records, and the table. 14 seconds. In Pi, a single-record lookup through
pi-mcp-adapter took 6 seconds.
- We checked
- All five PMIDs, DOIs, titles, first authors, and years matched PubMed's own records. Nothing was invented.
- Watch for
- Relevance was mixed: one result was a position paper, and two were only partly about rural care. Accurate metadata doesn't mean the paper belongs in your review.
Table 1 of the Census Bureau report Health Insurance Coverage in the United States: 2023 (P60-284, 28 pages), extracted locally with the liteparse skill.
Prompt
Use the liteparse skill to extract Table 1 from sources/p60-284.pdf into outputs/table1.csv. Use one row per coverage type with columns: coverage_type, level, number_2022, moe_2022, number_2023, moe_2023, all in thousands as printed. Remove footnote markers from the labels. Record the PDF page and printed page number in outputs/table1-notes.md, along with the table's units and footnotes. Do the extraction with a script in scripts/.
- What it did
- Parsed the PDF on the Mac, found the table, and wrote a script that produced the CSV. About 3 minutes.
- We checked
- All 48 values (12 rows × 4 columns) matched the printed table. Footnote markers were removed from labels such as "Employment-based²", and both page numbers were right (PDF page 8, printed page 2).
- Watch for
- The footnotes in the notes file were garbled: text from the next column was mixed in. The run also used about 5 million input tokens, because the agent re-read parser output many times.
The 15-page paper "Attention Is All You Need" converted to Markdown with the MarkItDown MCP server, then questioned about its results.
- What it did
- Converted the PDF, listed its numbered headings two levels deep (19 in all; it skipped the third-level 3.2.1–3.2.3), and reported both BLEU scores from the results table, quoting the table row. 81 seconds.
- We checked
- The scores (28.4 and 41.8) and the listed headings match the paper.
- Watch for
- Some lines lost the spaces between words ("EncoderandDecoderStacks"). The session used about 625,000 input tokens, because the full text went into the context more than once.
Telehealth for adult mental health care in the rural United States, 2020–2025: searches in PubMed and OpenAlex, deduplication, title-and-abstract screening, PRISMA counts, and a BibTeX file. Five prompts in one workspace.
Steps of the practice review with time, what happened, and what we found
| Step | Time | Result | What we found |
| Search | 53 s | 80 records, both searches logged | All 40 PubMed abstracts were empty (a parser bug), and the date limit used entry date instead of publication date. |
| Check | 3.7 min | Both bugs found and fixed | PubMed count went from 55 to 58, matching our own query. Its fix for one bad year was a special case, not a general rule. |
| Screen (first try) | 69 s | 19 include, 6 exclude, 54 unsure | It wrote keyword rules instead of reading: app matched "approach". Our house rules were too vague. We fixed them. |
| Screen (again) | 2.4 min | 13 include, 59 exclude, 7 unsure | 8 of 13 includes were right. 18 of 20 sampled excludes were defensible; one gave a false reason. |
| Bibliography | 39 s | 13 BibTeX entries | Every DOI resolved and every title matched Crossref. Three used the online-first year. |
- Watch for
- Twice the agent said "done" when the output was wrong. The check prompt in Kit 13, step 4, caught the first problem; reading the screening file caught the second.
A small Python package that scores Likert items. Two tests failed, and skipped answers (None) crashed the mean. We used the Plan agent first, then Build.
Prompt · Plan agent
Two tests in tests/test_scales.py fail. Find the cause. Also, survey respondents sometimes skip items; those arrive as None and scale_mean should ignore them. Propose a plan: the code changes and the new tests. Don't edit anything yet.
- What it did
- Plan (25 s): found that
reverse_score ignored the scale minimum, and proposed skipping None, dividing by answered items, and three new tests. Build (29 s): made the changes, and all 6 tests passed. It didn't commit, as its house rules say.
- We checked
- The diff is correct. We tested extra cases (0–4 and 1–7 scales, all items skipped); all were right.
- Watch for
- Reading the diff found a gap the agent missed:
scale_mean can't pass a scale minimum to reverse_score, so a reverse-scored item on a 0–4 scale is still wrong. Passing tests aren't the same as correct code.