Skip to content
View Selrach84's full-sized avatar
💭
Weh?
💭
Weh?

Block or report Selrach84

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Selrach84/README.md

Charles Denzel Segovia

AI Systems Engineer — agentic infrastructure, RAG, LLM evaluation. Founder of MagnifAI · Lapu-Lapu City, Cebu, PH.

I build and operate production LLM-agent infrastructure: multi-VM agent fleets driving dozens of integration CLIs, with a verification-first practice — every shipped claim carries evidence, every incident a documented root cause.

18 years of operations leadership underneath. That is the part that changes how I build: I have been the person woken up by the system, so I design for the 3 a.m. case rather than the demo case.


What I actually do

Agentic infrastructure Multi-VM agent fleets, ~70 integration CLIs, API-first with browser-automation fallback
RAG / retrieval Hybrid semantic + keyword search over personal and client knowledge bases, citation enforcement, groundedness scoring
Auth & session architecture Per-tenant session isolation, cookie-scope correctness, SAML/IdP boundary analysis
LLM evaluation Falsification-first debugging, multi-model adversarial panels, honest-uncertainty monitoring
Reverse-engineering SaaS Building CLI access to services that ship no API — session-backed, headless-browser driven

Featured work

Fixed fleet-wide auth kick-outs across 7 independently-deployed portals that each ran their own database but shared one cookie domain — an SSO topology that could not work.

Found five simultaneous defects by live probing rather than trusting documentation: a 7-day session expiry, a cookie scoped to a shared parent domain, a half-auth state where get-session returned null while an authenticated route returned HTTP 200, a 404 on the client route, and a single-entry front-door map in a multi-tenant deployment.

Shipped a 4-file patch per box, verified on every one.

The transferable lesson: a shared cookie domain is not a shared session.

Kept an automated agent logged into an MLS that provably cannot be re-authenticated programmatically — by never letting the session lapse.

I falsified four competing re-authentication hypotheses before writing any solution (page navigation does not extend the auth cookie — measured 4.1h → 4.1h; the IdP session does not silently re-issue; the login form does not auto-submit — verified in the DOM, not inferred; the portal SAML chain is a separate session).

The negative result inverted the design: from "re-authenticate" to "never let it lapse" — a cookie-cache bridge where the remote VM is fully autonomous.

The transferable lesson: absence of evidence that an approach works is not evidence that it cannot. Test to a structural negative.

API first, browser second, never the reverse — a pattern for giving an agent reliable access to ~70 SaaS integrations, plus the fleet-verification doctrine that came out of running it.

Encodes the three failure modes that all look healthy from the outside: false redundancy (a dual binary whose two lanes are one implementation), probe-contract failure reported as product failure (exit 2 means bad flag, not service down), and inventory drift (an archived inventory missed 7 of 11 live lanes; a same-day audit missed a batch deployed nine hours later).

An unavailable probe is UNKNOWN, not UNHEALTHY. Staleness is a trigger to PROBE, never to ACT.


Selected results

All from contract AI-systems work for a US real-estate AI platform (Jul–Sep 2026). Client-identifying details removed; every figure below was verified against live system output at the time.

  • Falsified four hypotheses before building a self-healing MLS keepalive — disproof was the deliverable that produced the design.
  • Root-caused a compounding nightly embedding failure across 167,274 chunks. Rejected the "dead worker" hypothesis; traced the real mechanism to a nightly regeneration that rewrites all pages at one instant. Drove the backlog 18 → 0 plus 4 chunks stranded in a previous model's vector space, and shipped a guard that re-measures and fails loudly.
  • Found live sales pipeline misreported as closed by auditing recycled CRM stage metadata — stage internal IDs no longer matched their labels and isClosed was wrong on the first four stages. Shipped label-based classification across all tooling. Framed as a discovery, not revenue.
  • Fixed fleet-wide portal auth across 7 production boxes, verified live on every one (see above).
  • Drove a fleet monitoring digest from 6 FAIL to 0 FAIL on evidence, not restarts. Triaged each failure to a distinct root cause and found that two of the six were monitoring bugs, not service faults — the naive fix would have restarted two healthy services and caused the outage it claimed to prevent.
  • Eliminated a monitoring false positive firing 18 of every 24 hours — a 6-hour threshold hardcoded against a 24-hour cycle, which was also triggering redundant compute.
  • Root-caused a metrics collector that had never been scheduled at all, serving data frozen for 43 days — surfacing a 5.24% opt-out rate against a 2% compliance limit that had been invisible the entire time.

How I work

Falsification before construction. The cheapest experiment is the one that kills an approach. Four disproved hypotheses cost less than one broken build.

Evidence over assertion. "It responds" is not connected; "HTTP 200" is not connected. A lane is connected when it returns real business data on five independent criteria — and a monitor that cannot say "I don't know" will manufacture failures.

Root cause before change. Correlation is not diagnosis. When a fix is proposed, the question is whether it changes the behaviour of the specific line where the symptom manifests.

Honest uncertainty. Where something was not verified, I say so. A verification record listing only green rows is a marketing document.


Stack

Python · TypeScript / Next.js · Go · SQL (Postgres, SQLite) · SQLite-vec · Docker · Linux · Chrome DevTools Protocol · MCP · OAuth / SAML / session architecture · Hermes · Claude Code · LLM evaluation harnesses


Elsewhere

Open to AI systems engineering and agentic-infrastructure work.

Popular repositories Loading

  1. obsidian-vault-rag obsidian-vault-rag Public

    Zero-dependency, heading-aware, 6-signal RAG for Obsidian — BM25 + recency + graph + fuzzy + tag + entity. Sub-15ms queries. MCP-native.

    Python 1

  2. headroom-rag-stack headroom-rag-stack Public

    Reproducible LLM token-optimization stack: Headroom proxy + RAG funnel + quality A/B proof

    Python

  3. MagnifAI-Showcase MagnifAI-Showcase Public

    A folder of various successful projects

    Python

  4. magnifai-ai-os magnifai-ai-os Public

    Python

  5. rag-chatbot rag-chatbot Public

    Chat with your Obsidian vault — hybrid RAG retrieval + DeepSeek LLM synthesis. Web UI with SSE streaming.

    HTML

  6. selrach84.github.io selrach84.github.io Public archive

    Charles Denzel Segovia — AI Automation Engineer portfolio

    HTML