Long context got cheap. Your prompt didn't.
Million-token windows are table stakes now. The bill didn't disappear — it moved somewhere most pipelines aren't looking.
AI Engineer · TitanCloud
Most of the work of an LLM pipeline happens upstream of the model. I write about the layers that filter, route, and validate before a single token is spent — and I build them for a living.
At TitanCloud I’m building a four-layer gatekeeper that cleans, filters, and routes documents before they reach a three-agent pipeline on Amazon Bedrock.
Million-token windows are table stakes now. The bill didn't disappear — it moved somewhere most pipelines aren't looking.
How a stack of cheap classifiers cuts a document-IDP bill by an order of magnitude without ever waking the model.
Routing, retrieval, and a small set of decisions about which questions a model should never see.
A three-agent IDP pipeline on Amazon Bedrock with a four-layer gatekeeper in front. Filters 92% of input before any LLM token is spent, routes the rest by intent, and closes the loop on human corrections.
A vision pipeline for live surface-defect detection on machined parts, published at NAMRC/MSEC 2025. Combined classical image features with a small CNN to keep inference under 30 ms on workshop hardware.
I’m an AI Engineer at TitanCloud. I finished an MS in Data Science at Gannon University in December 2025.
Before TitanCloud I shipped on-device computer vision at BitsKraft and ran the shared ML infrastructure for a research group of seven faculty. I care more about the boring layers of a system than the model at the bottom — because the model is rarely what’s broken.
I’m looking for a full-time AI / ML role starting mid-2026. F-1 OPT, open to relocation.
Best way to reach me is email. I read everything and reply to most things within a day or two.