DAVID RUSSELL

Programmer

David Russell outdoors

Model training

Fine-tuning the Story Copilot model

For Story Copilot, I trained a model to respond to players in a tabletop game. It needs to answer their questions, remember who knows what, and leave their decisions to them.

The latest adapter wrote much shorter replies, but did not clearly beat the base model on the separate test. The write-up covers the data, training setup, comparison and remaining failures.

Read the training results →
1,250reviewed training exchanges
27Bparameters in the base model
2 × 3090local training hardware

Applications and model tools I’m building. Each page explains the implementation, shows an example, and links the code and evaluation findings.

Hosted demo

Frontline

Draft maintenance reports from field notes, recordings and photos. Review the source for a finding, check applicable equipment guidance, and distinguish a completed repair from a suggested next step.

FastHTML · Groq · SQLite FTS5 / BM25

FRONTLINEAuthored report excerpt

Intermittent gateway outage

WORK

Reseated the DC connector. The spare supply was not installed.

CHECK

216 queued readings uploaded; reporting continued during a five-minute check.

FOLLOW-UP

Verify the supply rating. The photographed label is unreadable.

The work order requested a replacement; the visit note says it did not happen.

macOS application

Over The Shoulder

Get help while working in an editor, cloud console or browser terminal. The app reads the screen and listens to the conversation, then proposes code, explanations or diagrams. You operate the tools and apply the changes.

Python / AppKit · Screen and audio processing

OVER THE SHOULDERAuthored task excerpt

Changing a webhook retry policy

SCREEN

A worker retries responses with status >= 500.

AUDIO

“Handle 429 too. Honor Retry-After, but stop if it exceeds ten seconds.”

PROPOSAL

Preserve replay safety and retry limits. A long server delay leaves the item queued.

The walkthrough shows the code change and the cases it must handle.

Local web application

Career Workbench

Work through career research and resume drafts with an agent that can ask follow-up questions and use saved evidence. Correct a source account and see which claims and profiles need another review.

FastHTML · Codex / MCP · SQLite · Typst

CAREER WORKBENCHFictional account

Who changed the returns process?

ACCOUNT

“I built the intake form and tracker across three branches.”

CORRECTION

“Finance also added an approver. The timing figures were estimates.”

DRAFT

Describe the process and coordination work; omit the unsupported speed claim.

Local GPU training and serving

Qwen training tools

Prepare conversational training data, train task-specific LoRAs, and compare them with the base model. I used these tools to train Story Copilot models on two RTX 3090s; the repos support running the same process with your own data.

PyTorch · QLoRA / FSDP2 · llama.cpp

QWEN TRAININGRecorded test result

Base model vs. prose adapter

BASE

27 pass · 8 fail · 5 uncertain

ADAPTER

30 pass · 9 fail · 1 uncertain

PREFERENCE

19 for each model, with 2 ties.

40 cases, one response per model per case. Assistant judgments; no human calibration.

Local web application

Story Copilot

Suggest the game facilitator’s next reply using the conversation, scenario and character sheets. It can look up an earlier exchange or a supplied rule before answering. Suggestions stay private and do not become recorded events.

Local models · Evidence tools · Conversation memory

STORY COPILOTAuthored scene excerpt

At the harbor signal station

INEZ

Questions Ada about a ferry departure. Keeps a private note to herself.

BRAM

Checks the view of the workshop from the doorway; stays outside.

REPLY

Give Ada an answer and Bram an observation, without disclosing the note or moving him inside.

GPU kernels and compiler tools

ML Compiler Lab

Compile model graphs and compare their GPU execution. CuTe kernels, a persistent tensor-core megakernel, and a C++ IREE scheduling pass make layouts, launch costs and intermediate memory use inspectable.

CuTe DSL · CUDA Graphs · PyTorch · IREE / MLIR · C++

ML COMPILER LABImplemented GPU dataflow

Two projections, one kernel launch

COMPUTE

A thread block owns sixteen independent rows and both tensor-core projections.

MEMORY

The FP16 hidden state stays in shared memory between stages.

CHECK

Compare against conventional launches, CUDA Graph replay and cuBLAS.

Compiler and C++ simulator

Tile Accelerator

Compile a neural-network block into tile transfers, matrix instructions and dependencies. Execute its binary, check the output, and study how scratch capacity and DMA overlap change an explicit hardware model.

Python · C++ · Binary ISA · IREE VM bridge

TILE ACCELERATORBinary execution and cost model

Make the transfers explicit

COMPILE

Load tiles, accumulate the matrix product, apply the epilogue and store the result.

EXECUTE

An independent C++ engine checks the generated binary and produces an output.

MODEL

Vary scratch memory, compute rate and DMA overlap; keep estimates labelled.