AI engineer. I build LLM systems that are safe to put in front of real users.
Grounded RAG · agent guardrails · human-in-the-loop · evaluation you can fail a build on
Most LLM demos work. Most LLM products break the first time someone asks a question the knowledge base cannot answer — and the model answers anyway, fluently and wrongly.
Closing that gap is the work. I design and ship retrieval and agent systems where the unglamorous parts are in place from day one:
Grounded, or honestly silent. Hybrid retrieval (pgvector similarity + PostgreSQL full-text, fused with Reciprocal Rank Fusion), reranking, evidence-sufficiency checks, contradiction resolution by source authority, and validation of every citation, number, date and entity before a user sees a word. No evidence, no answer — the system says so and offers a human instead of guessing.
Governance as a deployment condition, not a retrofit. Tool calls classified by risk, human approval gates in front of irreversible actions, unknown tools failing closed to high risk, per-call cost tracking, and an audit trail behind every decision.
Production shape from the first commit. Multi-tenancy enforced in SQL, role-based access control, versioned document ingestion with atomic activation, prompt-injection defences, structured logs and per-stage traces, migrations, Docker, CI.
Measured, not vibed. Golden-dataset evaluation with enforced thresholds for groundedness, citation correctness, abstention accuracy and hallucination rate — running in CI, allowed to fail the build.
| Project | What it is |
|---|---|
| enterprise-rag-platform | A multi-tenant enterprise knowledge search product. Four retrieval legs — analyzed BM25, exact-identifier, dense kNN and whole-section — issued as a single OpenSearch _msearch, fused with Reciprocal Rank Fusion in application code, then reranked by a cross-encoder. Permission filters live inside the query, including the kNN filter, so narrow-ACL users lose no recall and document existence never leaks through result counts. Zero-downtime index rebuilds behind a shadow-evaluation gate a generation cannot go live without passing. SSO (Entra, Google, SAML via a Keycloak broker), SCIM deprovisioning, and a vendor-support plane where reading customer content requires a time-boxed grant the customer approved. ~1,100 tests, and an ablation table in CI that makes every stage justify its latency. |
| AI-Agent-Customer-Support-RAG | A support agent that answers only from an approved knowledge base, cites the exact evidence, and abstains or hands off to a human when the evidence is missing, weak or contradictory. FastAPI + pgvector hybrid RAG, React support console, multi-tenant, evaluated against a golden dataset in CI. |
| agent-guardrail | A framework-agnostic safety layer between an agent and its tools: risk classification, human approval for irreversible calls, Bedrock Guardrails content/PII checks, cost tracking and a full audit log — in ~250 lines of FastAPI. |
| rag-assistant | A fully local RAG assistant over your own PDFs, powered by Claude. |
Languages Python · SQL
AI Anthropic Claude · OpenAI · Azure OpenAI · Amazon Bedrock · Ollama · RAG · agents & tool use · evaluation harnesses
Backend FastAPI · Pydantic · SQLAlchemy (async) · Alembic · PostgreSQL + pgvector · Neo4j
Cloud & ops AWS (Lambda, ECS, S3, Bedrock, API Gateway) · Docker · GitHub Actions · structured logging
A model that is confidently wrong once costs more trust than a hundred correct answers earn.
So my systems are allowed to say "I don't know" — and are built so that they actually do.


