Senior full-stack AI engineer with 10+ years of shipping products. I'm at Comet ML working on OPIK, an open-source platform for debugging, evaluating and monitoring LLM applications. Its Python SDK gets about 1.7M downloads a month.
On OPIK I own the frontend. I built the TypeScript SDK from scratch and the Ollie AI assistant end to end, from the streaming UI and tool calls to its Python side. I also own the hosted MCP server, including its OAuth in Java. Most days that means TypeScript, React, Node, Python and PostgreSQL, with OpenTelemetry traces running into the millions of rows.
Barcelona · yaroslavboiko.com · LinkedIn · Email
| Project | What I did | Stars |
|---|---|---|
| OPIK | Core contributor to Comet ML's LLM evaluation and observability platform, with 351 public PRs across the frontend, the TypeScript SDK, the Ollie assistant and the Java and Python backends. | |
| notion-mcp-server | MCP server for Notion pages and databases. Dozens of operations behind two public tools, so it doesn't flood the agent's context. | |
| replicate-flux-mcp | Small MCP server for Replicate Flux image generation with inputs and outputs you can inspect. | |
| OPIK MCP | Comet ML's MCP server that connects OPIK prompts, projects, traces and metrics to coding assistants. I own the hosted version and its OAuth. | |
| yaroslavboiko.com | My site and blog. Astro on Cloudflare Workers, plus a Three.js scene that earns its bytes. |
My two MCP packages get around 3.5K npm downloads a week.
- AI agents in production. Streaming, sessions, tool calling, MCP clients, guardrails and the harness around them, built for Ollie.
- MCP servers. Tool surfaces that stay small, errors that tell the agent how to fix its next call, and OAuth for hosted servers.
- TypeScript SDKs and APIs. Typed contracts, compatibility tradeoffs, and examples people can copy.
- Evals and LLM cost. Offline evals, prompt optimization, and tracking what each LLM call costs.
- Performance. Trace views over millions of rows, and a week spent on performance during a customer POC that ended with a signed deal.
- MCP Tool Design for AI Agents, Not API Endpoints: why I expose one execute tool and one describe tool instead of an endpoint per operation.
- Stop Using Claude Code on Defaults: five settings I changed in
~/.claude/settings.jsonto save tokens and stop approvinglsfor the 400th time. - Agentic UX Primitives: streaming, HITL gates, reasoning traces and confidence indicators, the frontend patterns behind Cursor and Claude.
- Context Engineering Ate Prompt Engineering: why structured context beats clever prompts, and how much of it to leave out.
Full archive: yaroslavboiko.com/blog
- I write the spec and the tradeoffs down before building anything clever.
- If behavior crosses product boundaries, it gets a typed SDK and examples, not another one-off adapter.
- Agent features ship with traces and offline evals. "Looks good in chat" doesn't count as testing.
- Anything that should feel alive streams. Actions that are expensive to get wrong get a human approval step.
- Agents get a tight working set of context, and prompts live in the repo like any other code.
- TypeScript end to end for product work, Python where eval or ML tooling makes it the better fit.





