Visitar URL original
raihan-js (Raihan) · GitHub
Skip to content
View raihan-js's full-sized avatar
🔥
Working from home
🔥
Working from home

Block or report raihan-js

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
raihan-js/README.md

Hey there, I'm Raihan 👋

AI/ML Engineer · LLMOps · Evaluation · Retrieval · From Bangladesh

I build the measurement layer for LLM systems: release gates for quantised models, label-free monitoring, judge audits and grounding checks, each reported with per-item data and paired statistics.

Portfolio Hugging Face dev.to LinkedIn Email


What I'm working on

  • 🧪 Twelve open research projects (below), each with code, tests and data on GitHub and Hugging Face. Nulls and negative results are reported as nulls, and every number carries its scope.
  • 🏛️ Founding engineer and AI/ML lead at VETR Proposal (contract): an AI proposal platform for federal contractors. I trained FedProc-180M, F1 0.800 vs 0.804 for Claude Haiku 4.5 on FAR-clause extraction (FedProc-Bench test set), with 13.8% vs 32.1% hallucinated clauses.
  • 🇯🇵 Japanese-language evaluation and retrieval: JaCite-Bench, tiny-bilingual-retriever and Invoice-Check JP below.
  • 🧱 Small models from scratch on one GPU: the ORCH code models and the Vocab Tax study.

Research projects

Every result below cites its dataset, sample size and hardware in the project's README. Most are measured on synthetic or small data and say so there.

Project What it does Key result (scope in the README) Links
FlipGate Release gate for quantised LLMs: counts per-item right→wrong flips against a measured noise floor Qwen2.5-3B, GSM8K-1000: AWQ and GPTQ each broke 91 correct answers (p = 0.0050 and 0.0028) while accuracy moved 3.5 to 3.7 points; the gate fails both GH · HF · 📝
ShiftWatch Label-free accuracy estimation under data shift, with a FastAPI + Prometheus sidecar No estimator wins everywhere (Banking77, CLINC150, ModernBERT-base); DoC and CBPE never detect a drop GH · HF · 📝
OracleBench Audits small LLM judges against deterministic oracles A 3B judge falsely accepts 11.0% of wrong answers, a 0.5B one 40.6% (1,760 items); checker-first harness: 0 errors by construction vs 412 for judge-only, 17.6× fewer judge calls GH · HF
JaCite-Bench Do LLMs invent Japanese law articles? Registry of 11 laws, 6,913 articles, 600 questions llm-jp-3-1.8b invents 4.05% of Japanese-language citations vs 1.09% in English (3.7×) GH · HF
GraphProof-QA Proof-carrying multi-hop QA with a 1.5B model 34.2% direct vs 93.3% DSL vs 96.8% constrained on 6,000 MetaQA questions; entities renamed to unseen strings: 81.4% vs 5.8% GH · HF
FedProc-Constrained A registry grammar for FAR clause numbers Fabricated clauses 57/60 free vs 0/60 with the grammar (Qwen2.5-1.5B); with an abstain option the model also refused real clauses GH · HF
tiny-bilingual-retriever bge-m3 (568M) distilled into a 36.7M English-Japanese encoder 71% of the teacher's EN-JA nDCG@10 (synthetic eval) with 15× fewer parameters and a 4× smaller index; the public ruri-v3-30m scores higher GH · HF
Roofline-First Decoding Fused W4A16 Triton GEMV for batch-1 decoding, written after computing the bandwidth ceiling RTX 3060: 25.5 tok/s vs 25.7 for bitsandbytes NF4, with lower probe perplexity; measured copy bandwidth 323.9 GB/s GH · HF
Vocab Tax Compute-matched vocabulary study, 16 decoders trained from scratch on TypeScript/JavaScript 8k to 16k vocabularies win at every size; an 8.0M model with an 8k vocabulary beats a 33.5M one with 2k (1.168 vs 1.436 bits/byte); seed noise up to 0.13 GH · HF
DemoDoctor What do bad robot demonstrations cost a policy? A stall detector reaches F1 0.867 on injected faults; 12 ACT policies on PushT show no detectable effect of corrupted data (grid underpowered) GH
Invoice-Check JP A fine-tuned VLM reads Japanese qualified invoices (適格請求書); check digit, registry lookup and tax arithmetic gate auto-approval QLoRA Qwen2.5-VL-3B on synthetic invoices: 83.0% exact; the checks auto-approve 86.8% with 0 of 461 wrong in a checkable field. Adapter is non-commercial GH · HF dataset · HF adapter
Keiri-Agent A LangGraph back-office agent for Japanese invoices: verification, PO matching, human review, PostgreSQL checkpoints, LangSmith evaluation, exact CI gate Pre-registered on 300 synthetic invoices: the checks cut unsafe auto-approvals from 17.0% to 4.0%; PO matching mostly rerouted rather than detected. Survives SIGKILL GH

More write-ups, charts and the full list: raihan-js.github.io.


Models I've trained

All published on 🤗 Hugging Face, with configs and tokenizers. Built and published, not deployed.

Model What it is Hardware
ORCH Next.js 3B 3B decoder-only LLaMA-style model trained from scratch for Next.js code generation (data-limited) one rented A40 48GB
ORCH-7B QLoRA fine-tune of DeepSeek Coder 6.7B Instruct (5,238 steps) one A100, 43 h
ORCH Fusion 272M model trained from scratch, with a custom 2,103-token vocabulary one RTX 3060 12GB
ORCH Next.js 350M v2 287M model trained from scratch with a 16k vocabulary one RTX 3060 12GB
FedProc-180M ModernBERT-base, 4 task heads, FAR-clause extraction see the model card

Earlier work, past role. At ClarioScope AI (CTO and lead AI engineer, 2024 to 2026; the company was sunset in 2026 after two pivots and the models were open-sourced) I built and published the ClarioScope SLM suite: a 184M intent classifier (91.2% vs 95.2% for GPT-4o on a held-out set, about 22× faster than Claude Haiku 4.5), a 125M PHI detector (18 HIPAA Safe Harbor categories) and a 125M insurance extractor (12 fields). Write-ups: the suite and the PHI detector.


Tech stack

AI & ML

PyTorch Hugging Face PEFT / QLoRA Triton vLLM LangGraph LangSmith

Backend & platform

Python FastAPI PostgreSQL Docker GitHub Actions AWS Laravel Node.js

Frontend

React Next.js TypeScript


Open source

  • orch-ai: Hugging Face org for the ORCH code-generation model family
  • clarioscope-ai: Hugging Face org for the ClarioScope models
  • Also engineer and maintainer of CommonRoom AI (a 15-mini-app collaboration app) and founder of ILMA Lang (a beginner programming language that transpiles to C)

📫 Get in touch: raihan@vetrproposal.com · Portfolio · LinkedIn

Pinned Loading

  1. 33-js-concepts 33-js-concepts Public

    Forked from leonardomso/33-js-concepts

    📜 33 JavaScript concepts every developer should know.

    JavaScript

  2. GPT-NEO-1.3B GPT-NEO-1.3B Public

    Python

  3. docs docs Public

    Forked from laravel/docs

    The Laravel documentation.

  4. yolov5-onnx-acnedataset yolov5-onnx-acnedataset Public

    HTML

  5. LLPhant/LLPhant LLPhant/LLPhant Public

    LLPhant - A comprehensive PHP Generative AI Framework using OpenAI GPT 4. Inspired by Langchain

    PHP 1.7k 173