projects.

Work

Computer Use Field Notes
Projects · Work ↗

Computer Use Field Notes

Eval runs to improve computer-use UX — OSWorld experiments on smaller screenshots, running task notes, and batched GUI actions, with every trace public.

DeepSeek wins IMO Gold on 12 cents
Projects · Work ↗

DeepSeek wins IMO Gold on 12 cents

DeepSeek V4 Flash scored 30/42 on IMO 2026 in Cline, clearing the gold cutoff for $0.12 — about 140 times cheaper than Claude Fable 5.

The Math of LLM Inference
Projects · Work ↗

The Math of LLM Inference

How to save millions by self-hosting LLMs — the math of inference and the real dollars, grounded in production traffic.

LLM Visualization
Projects · Work ↗

LLM Visualization

A GPT drawn at full resolution — all 85,728 parameters, every activation, and a guided walkthrough from tokens to output probabilities. Runs the model live in your browser.

Open-sourcing evals for open-weight agents
Projects · Work ↗

Open-sourcing evals for open-weight agents

The Hill Climber's Checklist for evaluating open-weight agents, with scores, tradeoffs, and public traces.

Public Speaking
Projects · Work ↗

Public Speaking

Talks and conference videos on AI agents, evals, and machine learning

Recursive Self Improvement for Coding Agents
Projects · Work ↗

Recursive Self Improvement for Coding Agents

How one prompt turned a 17-hour autonomous hill climb into an 88.8% state-of-the-art Terminal-Bench 2.1 score.

A Practical Guide to Hill Climbing
Projects · Work ↗

A Practical Guide to Hill Climbing

Iterative improvement for AI coding agents

Ruby TensorFlow
Projects · Work ↗

Ruby TensorFlow

Porting TensorFlow to Ruby with the tensorflow.rb gem — the Ruby API, how Google Protobuf drives the graph internals, and Inception-v3 image recognition in Ruby.

Agentic AI writings
Projects · Work ↗

Agentic AI writings

Art