The same work, pointed at your problem
I help businesses run models privately, specialize them for one job, and prove they work before anyone commits.
Private AI deployment
For legal, medical, financial, and regulated retail work where sending data to someone else’s API isn’t an option. I pick the model, quantize it to fit what you own, set up serving, and hand it over with documentation your team can run without me.
What you get
- Model selection against your tasks
- Quantization sized to your hardware
- A serving setup your team can operate
- A written handoff
Start a conversation about private AI deployment
Domain fine-tuning and distillation
The same method as my research: take what much larger models do well in your domain and train it into a model that fits on one GPU. You get a specialist that does your job better than a generalist of the same size, and you own the weights.
What you get
- Data curation from your domain
- A trained, evaluated specialist
- Weights you own
- Before and after numbers on your tasks
Start a conversation about domain fine-tuning and distillation
Evaluation and model selection
Public leaderboards don’t measure your work. I build a harness from your real tasks, run the candidates through it the way I evaluate my own releases, and give you a clear recommendation with the numbers behind it, including where each model falls short.
What you get
- A harness built from your tasks
- Side by side results
- A clear recommendation
- The harness itself, to rerun later
Start a conversation about evaluation and model selection
AI integration into existing systems
I’ve spent years connecting point-of-sale systems, storefronts, and apps. That means a model can go where the work already happens instead of into another tab nobody opens.
What you get
- Integration into your current stack
- Guardrails and fallbacks
- Monitoring you can read
- Documentation
Start a conversation about aI integration into existing systems
How an engagement runs
Conversation
You tell me what you’re trying to do. I tell you honestly whether a model is the right tool.
Scoping
A short written plan: what gets built, how we’ll know it works, and what it needs from you.
Evaluation first
Before building anything, candidate approaches get tested on your real tasks.
Build
Deployment, training, or integration, with progress you can see along the way.
Handoff
Documentation, the evaluation harness, and everything needed to run it without me.
Why me
I've done this work on my own models and published it: Fusion and GLM-18B as first author, and training compute, curation, and evaluation for the Qwopus line. My own releases have been downloaded more than a million times.
Since 2020 I've also been ecommerce director and lead developer for a specialty outdoor retailer, and I build apps and point-of-sale integrations for other businesses. I know what it takes for a model to be useful inside a real operation, not just on a benchmark.
It's one person. You talk to the person doing the work.