DAVID RUSSELL
Programmer
Model training
Fine-tuning the Story Copilot model
For Story Copilot, I trained a model to respond to players in a tabletop game. It needs to answer their questions, remember who knows what, and leave their decisions to them.
The latest adapter wrote much shorter replies, but did not clearly beat the base model on the separate test. The write-up covers the data, training setup, comparison and remaining failures.
Read the training results →
1,250reviewed training exchanges
27Bparameters in the base model
2 × 3090local training hardware
Applications and model tools I’m building. Each page explains the implementation, shows an example, and links the code and evaluation findings.
Hosted demo
Draft maintenance reports from field notes, recordings and photos. Review the source for a finding, check applicable equipment guidance, and distinguish a completed repair from a suggested next step.
FastHTML · Groq · SQLite FTS5 / BM25
FRONTLINEAuthored report excerpt
Intermittent gateway outage
WORK Reseated the DC connector. The spare supply was not installed.
CHECK 216 queued readings uploaded; reporting continued during a five-minute check.
FOLLOW-UP Verify the supply rating. The photographed label is unreadable.
macOS application
Get help while working in an editor, cloud console or browser terminal. The app reads the screen and listens to the conversation, then proposes code, explanations or diagrams. You operate the tools and apply the changes.
Python / AppKit · Screen and audio processing
OVER THE SHOULDERAuthored task excerpt
Changing a webhook retry policy
SCREEN A worker retries responses with status >= 500.
AUDIO “Handle 429 too. Honor Retry-After, but stop if it exceeds ten seconds.”
PROPOSAL Preserve replay safety and retry limits. A long server delay leaves the item queued.
Local web application
Work through career research and resume drafts with an agent that can ask follow-up questions and use saved evidence. Correct a source account and see which claims and profiles need another review.
FastHTML · Codex / MCP · SQLite · Typst
CAREER WORKBENCHFictional account
Who changed the returns process?
ACCOUNT “I built the intake form and tracker across three branches.”
CORRECTION “Finance also added an approver. The timing figures were estimates.”
DRAFT Describe the process and coordination work; omit the unsupported speed claim.
Local GPU training and serving
Prepare conversational training data, train task-specific LoRAs, and compare them with the base model. I used these tools to train Story Copilot models on two RTX 3090s; the repos support running the same process with your own data.
PyTorch · QLoRA / FSDP2 · llama.cpp
QWEN TRAININGRecorded test result
Base model vs. prose adapter
BASE 27 pass · 8 fail · 5 uncertain
ADAPTER 30 pass · 9 fail · 1 uncertain
PREFERENCE 19 for each model, with 2 ties.
Local web application
Suggest the game facilitator’s next reply using the conversation, scenario and character sheets. It can look up an earlier exchange or a supplied rule before answering. Suggestions stay private and do not become recorded events.
Local models · Evidence tools · Conversation memory
STORY COPILOTAuthored scene excerpt
At the harbor signal station
INEZ Questions Ada about a ferry departure. Keeps a private note to herself.
BRAM Checks the view of the workshop from the doorway; stays outside.
REPLY Give Ada an answer and Bram an observation, without disclosing the note or moving him inside.
GPU kernels and compiler tools
Compile model graphs and compare their GPU execution. CuTe kernels, a persistent tensor-core megakernel, and a C++ IREE scheduling pass make layouts, launch costs and intermediate memory use inspectable.
CuTe DSL · CUDA Graphs · PyTorch · IREE / MLIR · C++
ML COMPILER LABImplemented GPU dataflow
Two projections, one kernel launch
COMPUTE A thread block owns sixteen independent rows and both tensor-core projections.
MEMORY The FP16 hidden state stays in shared memory between stages.
CHECK Compare against conventional launches, CUDA Graph replay and cuBLAS.
Compiler and C++ simulator
Compile a neural-network block into tile transfers, matrix instructions and dependencies. Execute its binary, check the output, and study how scratch capacity and DMA overlap change an explicit hardware model.
Python · C++ · Binary ISA · IREE VM bridge
TILE ACCELERATORBinary execution and cost model
Make the transfers explicit
COMPILE Load tiles, accumulate the matrix product, apply the epilogue and store the result.
EXECUTE An independent C++ engine checks the generated binary and produces an output.
MODEL Vary scratch memory, compute rate and DMA overlap; keep estimates labelled.