AI/ML Engineer · LLMOps · Evaluation · Retrieval · From Bangladesh
I build the measurement layer for LLM systems: release gates for quantised models, label-free monitoring, judge audits and grounding checks, each reported with per-item data and paired statistics.
- 🧪 Twelve open research projects (below), each with code, tests and data on GitHub and Hugging Face. Nulls and negative results are reported as nulls, and every number carries its scope.
- 🏛️ Founding engineer and AI/ML lead at VETR Proposal (contract): an AI proposal platform for federal contractors. I trained FedProc-180M, F1 0.800 vs 0.804 for Claude Haiku 4.5 on FAR-clause extraction (FedProc-Bench test set), with 13.8% vs 32.1% hallucinated clauses.
- 🇯🇵 Japanese-language evaluation and retrieval: JaCite-Bench, tiny-bilingual-retriever and Invoice-Check JP below.
- 🧱 Small models from scratch on one GPU: the ORCH code models and the Vocab Tax study.
Every result below cites its dataset, sample size and hardware in the project's README. Most are measured on synthetic or small data and say so there.
| Project | What it does | Key result (scope in the README) | Links |
|---|---|---|---|
| FlipGate | Release gate for quantised LLMs: counts per-item right→wrong flips against a measured noise floor | Qwen2.5-3B, GSM8K-1000: AWQ and GPTQ each broke 91 correct answers (p = 0.0050 and 0.0028) while accuracy moved 3.5 to 3.7 points; the gate fails both | GH · HF · 📝 |
| ShiftWatch | Label-free accuracy estimation under data shift, with a FastAPI + Prometheus sidecar | No estimator wins everywhere (Banking77, CLINC150, ModernBERT-base); DoC and CBPE never detect a drop | GH · HF · 📝 |
| OracleBench | Audits small LLM judges against deterministic oracles | A 3B judge falsely accepts 11.0% of wrong answers, a 0.5B one 40.6% (1,760 items); checker-first harness: 0 errors by construction vs 412 for judge-only, 17.6× fewer judge calls | GH · HF |
| JaCite-Bench | Do LLMs invent Japanese law articles? Registry of 11 laws, 6,913 articles, 600 questions | llm-jp-3-1.8b invents 4.05% of Japanese-language citations vs 1.09% in English (3.7×) | GH · HF |
| GraphProof-QA | Proof-carrying multi-hop QA with a 1.5B model | 34.2% direct vs 93.3% DSL vs 96.8% constrained on 6,000 MetaQA questions; entities renamed to unseen strings: 81.4% vs 5.8% | GH · HF |
| FedProc-Constrained | A registry grammar for FAR clause numbers | Fabricated clauses 57/60 free vs 0/60 with the grammar (Qwen2.5-1.5B); with an abstain option the model also refused real clauses | GH · HF |
| tiny-bilingual-retriever | bge-m3 (568M) distilled into a 36.7M English-Japanese encoder | 71% of the teacher's EN-JA nDCG@10 (synthetic eval) with 15× fewer parameters and a 4× smaller index; the public ruri-v3-30m scores higher | GH · HF |
| Roofline-First Decoding | Fused W4A16 Triton GEMV for batch-1 decoding, written after computing the bandwidth ceiling | RTX 3060: 25.5 tok/s vs 25.7 for bitsandbytes NF4, with lower probe perplexity; measured copy bandwidth 323.9 GB/s | GH · HF |
| Vocab Tax | Compute-matched vocabulary study, 16 decoders trained from scratch on TypeScript/JavaScript | 8k to 16k vocabularies win at every size; an 8.0M model with an 8k vocabulary beats a 33.5M one with 2k (1.168 vs 1.436 bits/byte); seed noise up to 0.13 | GH · HF |
| DemoDoctor | What do bad robot demonstrations cost a policy? | A stall detector reaches F1 0.867 on injected faults; 12 ACT policies on PushT show no detectable effect of corrupted data (grid underpowered) | GH |
| Invoice-Check JP | A fine-tuned VLM reads Japanese qualified invoices (適格請求書); check digit, registry lookup and tax arithmetic gate auto-approval | QLoRA Qwen2.5-VL-3B on synthetic invoices: 83.0% exact; the checks auto-approve 86.8% with 0 of 461 wrong in a checkable field. Adapter is non-commercial | GH · HF dataset · HF adapter |
| Keiri-Agent | A LangGraph back-office agent for Japanese invoices: verification, PO matching, human review, PostgreSQL checkpoints, LangSmith evaluation, exact CI gate | Pre-registered on 300 synthetic invoices: the checks cut unsafe auto-approvals from 17.0% to 4.0%; PO matching mostly rerouted rather than detected. Survives SIGKILL | GH |
More write-ups, charts and the full list: raihan-js.github.io.
All published on 🤗 Hugging Face, with configs and tokenizers. Built and published, not deployed.
| Model | What it is | Hardware |
|---|---|---|
| ORCH Next.js 3B | 3B decoder-only LLaMA-style model trained from scratch for Next.js code generation (data-limited) | one rented A40 48GB |
| ORCH-7B | QLoRA fine-tune of DeepSeek Coder 6.7B Instruct (5,238 steps) | one A100, 43 h |
| ORCH Fusion | 272M model trained from scratch, with a custom 2,103-token vocabulary | one RTX 3060 12GB |
| ORCH Next.js 350M v2 | 287M model trained from scratch with a 16k vocabulary | one RTX 3060 12GB |
| FedProc-180M | ModernBERT-base, 4 task heads, FAR-clause extraction | see the model card |
Earlier work, past role. At ClarioScope AI (CTO and lead AI engineer, 2024 to 2026; the company was sunset in 2026 after two pivots and the models were open-sourced) I built and published the ClarioScope SLM suite: a 184M intent classifier (91.2% vs 95.2% for GPT-4o on a held-out set, about 22× faster than Claude Haiku 4.5), a 125M PHI detector (18 HIPAA Safe Harbor categories) and a 125M insurance extractor (12 fields). Write-ups: the suite and the PHI detector.
- orch-ai: Hugging Face org for the ORCH code-generation model family
- clarioscope-ai: Hugging Face org for the ClarioScope models
- Also engineer and maintainer of CommonRoom AI (a 15-mini-app collaboration app) and founder of ILMA Lang (a beginner programming language that transpiles to C)
📫 Get in touch: raihan@vetrproposal.com · Portfolio · LinkedIn




