<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/"><channel><title>17th Street Labs Blog</title><description>Practical AI engineering lessons: LLM evals, agents, model costs, local models, and AI security.</description><link>https://17thstreetlabs.com/</link><language>en-us</language><item><title>New Models, 3 New Evals: OCR, Translation, and Security Verification Tests</title><link>https://17thstreetlabs.com/blog/holy-model-release-week/</link><guid isPermaLink="true">https://17thstreetlabs.com/blog/holy-model-release-week/</guid><description>GPT-6, Claude, Gemini, GLM, DeepSeek and Qwen on our handwriting OCR, translation and security-verifier tests—with the costs, ties and judge disagreements left in.</description><pubDate>Thu, 24 Sep 2026 19:32:48 GMT</pubDate><dc:creator>Dan Levy</dc:creator><category>Testing &amp; Evaluation</category></item><item><title>Small Models. Serious Work.</title><link>https://17thstreetlabs.com/blog/small-local-models/</link><guid isPermaLink="true">https://17thstreetlabs.com/blog/small-local-models/</guid><description>What small local models can do, where they struggle, and how to give them a fair test.</description><pubDate>Mon, 14 Sep 2026 22:07:01 GMT</pubDate><dc:creator>Marina Levy</dc:creator><dc:creator>Dan Levy</dc:creator><category>Agents</category></item><item><title>The Cheap Model Can Get Expensive.</title><link>https://17thstreetlabs.com/blog/cost-per-completed-task/</link><guid isPermaLink="true">https://17thstreetlabs.com/blog/cost-per-completed-task/</guid><description>The number worth watching is cost per completed task. Here’s how to measure it.</description><pubDate>Mon, 14 Sep 2026 22:07:01 GMT</pubDate><dc:creator>Marina Levy</dc:creator><dc:creator>Dan Levy</dc:creator><category>Costs and Optimization</category></item><item><title>Before You Buy the GPU.</title><link>https://17thstreetlabs.com/blog/renting-vs-buying-gpus/</link><guid isPermaLink="true">https://17thstreetlabs.com/blog/renting-vs-buying-gpus/</guid><description>Renting compute can buy you something more useful than hardware: room to change your mind.</description><pubDate>Mon, 14 Sep 2026 22:07:01 GMT</pubDate><dc:creator>Marina Levy</dc:creator><dc:creator>Dan Levy</dc:creator><category>Costs and Optimization</category></item><item><title>One Harness Does Not Fit Every Model.</title><link>https://17thstreetlabs.com/blog/one-harness-does-not-fit-every-model/</link><guid isPermaLink="true">https://17thstreetlabs.com/blog/one-harness-does-not-fit-every-model/</guid><description>The structure that helps a small model can get in a stronger model’s way.</description><pubDate>Mon, 14 Sep 2026 22:07:01 GMT</pubDate><dc:creator>Marina Levy</dc:creator><dc:creator>Dan Levy</dc:creator><category>Agents</category></item><item><title>The Attackers Are Getting Agents, Too.</title><link>https://17thstreetlabs.com/blog/continuous-security-testing/</link><guid isPermaLink="true">https://17thstreetlabs.com/blog/continuous-security-testing/</guid><description>Cheap automation changes the pace. Your security testing needs to keep up.</description><pubDate>Mon, 14 Sep 2026 22:07:01 GMT</pubDate><dc:creator>Marina Levy</dc:creator><dc:creator>Dan Levy</dc:creator><category>Security &amp; Privacy</category></item><item><title>Your AI Judge Needs to Be Judged, Too.</title><link>https://17thstreetlabs.com/blog/eval-driven-agent-development/</link><guid isPermaLink="true">https://17thstreetlabs.com/blog/eval-driven-agent-development/</guid><description>A practical approach to eval-driven development: test the journey, then test the grader.</description><pubDate>Mon, 14 Sep 2026 22:07:01 GMT</pubDate><dc:creator>Marina Levy</dc:creator><dc:creator>Dan Levy</dc:creator><category>Testing &amp; Evaluation</category></item><item><title>Your Tests Passed. Your Interface Didn’t.</title><link>https://17thstreetlabs.com/blog/agents-that-see-the-ui/</link><guid isPermaLink="true">https://17thstreetlabs.com/blog/agents-that-see-the-ui/</guid><description>Give an agent the browser recording, and it can look for what the final screenshot missed.</description><pubDate>Mon, 14 Sep 2026 22:07:01 GMT</pubDate><dc:creator>Marina Levy</dc:creator><dc:creator>Dan Levy</dc:creator><category>Testing &amp; Evaluation</category></item><item><title>Some Data Doesn’t Get to Leave.</title><link>https://17thstreetlabs.com/blog/local-ai-data-privacy/</link><guid isPermaLink="true">https://17thstreetlabs.com/blog/local-ai-data-privacy/</guid><description>Local AI can keep sensitive work inside your environment. Check the whole route.</description><pubDate>Mon, 14 Sep 2026 22:07:01 GMT</pubDate><dc:creator>Marina Levy</dc:creator><dc:creator>Dan Levy</dc:creator><category>Security &amp; Privacy</category></item><item><title>The Agent Is Working. Does Anyone Know What It’s Doing?</title><link>https://17thstreetlabs.com/blog/security-copilot-that-explains/</link><guid isPermaLink="true">https://17thstreetlabs.com/blog/security-copilot-that-explains/</guid><description>A security copilot should make the work easier to follow—and easier to question.</description><pubDate>Mon, 14 Sep 2026 22:07:01 GMT</pubDate><dc:creator>Marina Levy</dc:creator><dc:creator>Dan Levy</dc:creator><category>Security &amp; Privacy</category></item></channel></rss>