<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Amartya Gaur</title><description>Notes on agent orchestration, evals, and technical hiring.</description><link>https://amartya-gaur.com/</link><item><title>A pipeline that cannot invent your UI</title><link>https://amartya-gaur.com/blog/a-pipeline-that-cannot-invent-your-ui/</link><guid isPermaLink="true">https://amartya-gaur.com/blog/a-pipeline-that-cannot-invent-your-ui/</guid><description>Generated product videos show software that does not exist. Fixing that is not a prompting problem. It means making every frame cite the line of code it came from, and refusing to finish when it cannot.</description><pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate></item><item><title>A container per job, without a daemon</title><link>https://amartya-gaur.com/blog/a-container-per-job-without-a-daemon/</link><guid isPermaLink="true">https://amartya-gaur.com/blog/a-container-per-job-without-a-daemon/</guid><description>Building a bespoke image for every unit of work, from a package list rather than a Dockerfile, and why having no build step at all is the security property, not a limitation.</description><pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Taking the network away from a Cloud Run job</title><link>https://amartya-gaur.com/blog/no-network-inside-a-cloud-run-job/</link><guid isPermaLink="true">https://amartya-gaur.com/blog/no-network-inside-a-cloud-run-job/</guid><description>Running untrusted code with no egress, on managed infrastructure, without a VM you have to operate. The trick is one prefix, and the reason it looks broken at first is that loopback starts down.</description><pubDate>Sun, 16 Aug 2026 00:00:00 GMT</pubDate></item><item><title>What a routing benchmark cannot measure</title><link>https://amartya-gaur.com/blog/what-a-routing-benchmark-cannot-measure/</link><guid isPermaLink="true">https://amartya-gaur.com/blog/what-a-routing-benchmark-cannot-measure/</guid><description>I set out to prove a routing architecture and killed my own hypothesis on day one. What survived is a condition the standard dialogue corpora fail, including one where not a single test dialogue contains the decision being benchmarked.</description><pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Proving a challenge discriminates before anyone takes it</title><link>https://amartya-gaur.com/blog/proving-a-challenge-discriminates/</link><guid isPermaLink="true">https://amartya-gaur.com/blog/proving-a-challenge-discriminates/</guid><description>Every hunr challenge ships with two solutions: one an expert would write, one a plausible engineer would. If the hidden tests can&apos;t tell them apart, it never goes live.</description><pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Why a sticky orchestrator beat a smarter router</title><link>https://amartya-gaur.com/blog/a-sticky-orchestrator/</link><guid isPermaLink="true">https://amartya-gaur.com/blog/a-sticky-orchestrator/</guid><description>We kept trying to make the router smarter. What actually fixed support agents was refusing to re-route mid-conversation.</description><pubDate>Mon, 18 May 2026 00:00:00 GMT</pubDate></item><item><title>739 cases that can fail a build</title><link>https://amartya-gaur.com/blog/739-cases/</link><guid isPermaLink="true">https://amartya-gaur.com/blog/739-cases/</guid><description>An eval suite nobody can block a deploy with is a dashboard. Here&apos;s what it took to make ours a gate.</description><pubDate>Mon, 02 Mar 2026 00:00:00 GMT</pubDate></item></channel></rss>