AI Systems Engineer — agentic infrastructure, RAG, LLM evaluation. Founder of MagnifAI · Lapu-Lapu City, Cebu, PH.
I build and operate production LLM-agent infrastructure: multi-VM agent fleets driving dozens of integration CLIs, with a verification-first practice — every shipped claim carries evidence, every incident a documented root cause.
18 years of operations leadership underneath. That is the part that changes how I build: I have been the person woken up by the system, so I design for the 3 a.m. case rather than the demo case.
| Agentic infrastructure | Multi-VM agent fleets, ~70 integration CLIs, API-first with browser-automation fallback |
| RAG / retrieval | Hybrid semantic + keyword search over personal and client knowledge bases, citation enforcement, groundedness scoring |
| Auth & session architecture | Per-tenant session isolation, cookie-scope correctness, SAML/IdP boundary analysis |
| LLM evaluation | Falsification-first debugging, multi-model adversarial panels, honest-uncertainty monitoring |
| Reverse-engineering SaaS | Building CLI access to services that ship no API — session-backed, headless-browser driven |
Fixed fleet-wide auth kick-outs across 7 independently-deployed portals that each ran their own database but shared one cookie domain — an SSO topology that could not work.
Found five simultaneous defects by live probing rather than trusting
documentation: a 7-day session expiry, a cookie scoped to a shared parent domain,
a half-auth state where get-session returned null while an authenticated
route returned HTTP 200, a 404 on the client route, and a single-entry front-door
map in a multi-tenant deployment.
Shipped a 4-file patch per box, verified on every one.
The transferable lesson: a shared cookie domain is not a shared session.
Kept an automated agent logged into an MLS that provably cannot be re-authenticated programmatically — by never letting the session lapse.
I falsified four competing re-authentication hypotheses before writing any
solution (page navigation does not extend the auth cookie — measured 4.1h → 4.1h;
the IdP session does not silently re-issue; the login form does not auto-submit —
verified in the DOM, not inferred; the portal SAML chain is a separate session).
The negative result inverted the design: from "re-authenticate" to "never let it lapse" — a cookie-cache bridge where the remote VM is fully autonomous.
The transferable lesson: absence of evidence that an approach works is not evidence that it cannot. Test to a structural negative.
API first, browser second, never the reverse — a pattern for giving an agent reliable access to ~70 SaaS integrations, plus the fleet-verification doctrine that came out of running it.
Encodes the three failure modes that all look healthy from the outside:
false redundancy (a dual binary whose two lanes are one implementation),
probe-contract failure reported as product failure (exit 2 means bad flag, not
service down), and inventory drift (an archived inventory missed 7 of 11 live
lanes; a same-day audit missed a batch deployed nine hours later).
An unavailable probe is
UNKNOWN, notUNHEALTHY. Staleness is a trigger toPROBE, never toACT.
All from contract AI-systems work for a US real-estate AI platform (Jul–Sep 2026). Client-identifying details removed; every figure below was verified against live system output at the time.
- Falsified four hypotheses before building a self-healing MLS keepalive — disproof was the deliverable that produced the design.
- Root-caused a compounding nightly embedding failure across 167,274 chunks.
Rejected the "dead worker" hypothesis; traced the real mechanism to a nightly
regeneration that rewrites all pages at one instant. Drove the backlog
18 → 0plus 4 chunks stranded in a previous model's vector space, and shipped a guard that re-measures and fails loudly. - Found live sales pipeline misreported as closed by auditing recycled CRM
stage metadata — stage internal IDs no longer matched their labels and
isClosedwas wrong on the first four stages. Shipped label-based classification across all tooling. Framed as a discovery, not revenue. - Fixed fleet-wide portal auth across 7 production boxes, verified live on every one (see above).
- Drove a fleet monitoring digest from 6 FAIL to 0 FAIL on evidence, not restarts. Triaged each failure to a distinct root cause and found that two of the six were monitoring bugs, not service faults — the naive fix would have restarted two healthy services and caused the outage it claimed to prevent.
- Eliminated a monitoring false positive firing 18 of every 24 hours — a 6-hour threshold hardcoded against a 24-hour cycle, which was also triggering redundant compute.
- Root-caused a metrics collector that had never been scheduled at all,
serving data frozen for 43 days — surfacing a
5.24%opt-out rate against a2%compliance limit that had been invisible the entire time.
Falsification before construction. The cheapest experiment is the one that kills an approach. Four disproved hypotheses cost less than one broken build.
Evidence over assertion. "It responds" is not connected; "HTTP 200" is not connected. A lane is connected when it returns real business data on five independent criteria — and a monitor that cannot say "I don't know" will manufacture failures.
Root cause before change. Correlation is not diagnosis. When a fix is proposed, the question is whether it changes the behaviour of the specific line where the symptom manifests.
Honest uncertainty. Where something was not verified, I say so. A verification record listing only green rows is a marketing document.
Python · TypeScript / Next.js · Go · SQL (Postgres, SQLite) · SQLite-vec
· Docker · Linux · Chrome DevTools Protocol · MCP · OAuth / SAML / session
architecture · Hermes · Claude Code · LLM evaluation harnesses
- MagnifAI — AI automation: agentic systems, RAG, multi-agent pipelines
- Portfolio / writing — selrach84.github.io
- LinkedIn — charles-denzel-segovia
- Email — segoviacharles01@gmail.com
Open to AI systems engineering and agentic-infrastructure work.
