# From the team

Insights on AI agents, software reliability, and the future of production operations.

**What It Actually Takes for AI to Map Your Production Systems**  
An agent cannot diagnose incidents without understanding your system and how it's connected. Here is what that involves.  
Chenggang Wu·August 26, 2026·5 min read

**System Understanding Is the Whole Game for AI SREs**  
Momento CTO Daniela Miao on why an AI SRE can only be trusted with alerting decisions if it understands how a system is built.  
Peter Farago·August 18, 2026·5 min read

**Engineering Knowledge Replicants**  
Tribal knowledge isn't collective. It's fragmented and fragile, and incident response depends on it. Why documentation and DIY agents still fall short for AI SREs.  
Chenggang Wu·July 28, 2026

**Winter is Coming for AI Engineering**  
Every VP of Engineering I talk to right now is spending money on Claude Code or Cursor. Most of them can't tell me what they're getting for it.  
Chenggang Wu·July 15, 2026·6 min read

**Why You Should Definitely DIY Dev Tools with AI. Sometimes.**  
Engineering teams are building their own AI dev tools more than ever. After comparing notes with a lot of them, here's where DIY pays off and where it doesn't.  
Chenggang Wu·June 30, 2026·5 min read

**ClickHouse + Herald: AI DevOps intelligence on fast, cost-efficient telemetry**  
Herald, an AI DevOps agent in the House Mates program, puts the telemetry you keep in ClickStack to work: predicting incidents, delivering root cause in minutes, and answering questions about your systems. ClickHouse customers can get up to $25,000 in free Herald usage.  
Peter Farago·June 25, 2026

**No Runbooks, No Problem: Snorkel AI Gets Day-One Results with Herald**  
93% accuracy on engineering questions. Incidents resolved in minutes. No runbooks. No RCA documents. No Slack history.  
Kartik Mathur·June 4, 2026·4 min read

**Heralding the Future**  
We're excited to share that RunLLM is now Herald.  
Vikram Sreekanti·May 27, 2026·4 min read

**When Your AI-Powered RCA Spews Pages of Useless Text**  
When your AI-powered RCA tool floods Slack with hallucinated walls of text, the problem isn't the model. It's the missing data engineering underneath it.  
Chenggang Wu·May 19, 2026·7 min read

**Could Your AI-Generated Code Destroy Your Company?**  
When everyone can build software, someone still has to keep it running. A reliability leader helped me understand how engineering organizations are facing a new influx of code from all sides.  
Chenggang Wu·May 5, 2026

**The Code Nobody Read Is Already in Production**  
Ben Sigelman argues that AI-generated code is a reliability crisis in slow motion, and what it means for how we observe production systems.  
Peter Farago·April 29, 2026

**The Future of Software is Production**  
Ship every piece of code you write directly into production.  
Vikram Sreekanti·April 21, 2026

**Why LLM-Over-Logs Is the Wrong Abstraction.**  
Dumping logs into an LLM causes high variance and latency. Learn the data engineering approach for AI SRE that prioritizes signal over context.  
Vikram Sreekanti·April 14, 2026

**I Don’t Care if AI Wrote the Code. You Own It.**  
SREcon Chair Heinrich Hartmann on why the age of AI-assisted engineering demands a radical return to design rigor.  
Peter Farago·April 7, 2026

**The SDLC is Dead. Long Live the SDLC.**  
You cannot review your way out of the new glut of AI code. Winning teams will learn faster from what reaches production.  
Vikram Sreekanti·April 1, 2026

**The On-Call Problem AI Can Actually Solve**  
Heinrich Hartmann argues AI’s most valuable role isn’t autonomous remediation. It’s ensuring on-call engineers have the context they need to fix incidents fast.  
Peter Farago·March 10, 2026

**AI-Created Code Is Putting Us in Debt**  
The velocity trap is real. Here is the new engineering framework for surviving the age of AI-generated code.  
Peter Farago·October 21, 2025

**Can AI Spot Outages Faster Than Your Customers?**  
How AI shortens detection time and prevents trust-eroding surprises  
Peter Farago·October 14, 2025

**The End of SRE Tribal Knowledge**  
How AI turns expert intuition into operational infrastructure  
Peter Farago·October 7, 2025

**Never Let a Good Incident Go to Waste**  
How AI turns firefighting into continuous learning  
Peter Farago·October 1, 2025

**Why SREs Need an AI Teammate**  
AI that clears the path so on-call engineers move faster  
Peter Farago·September 25, 2025

**Respecting Control by Design**  
Principles for adding an AI teammate to incident response  
Peter Farago·September 17, 2025

**The Glass Box AI SRE**  
Why Transparency Wins in Incident Response  
Peter Farago·September 9, 2025

**More Needle, Less Haystack: Solving the AI SRE Trust Gap**  
How AI-assisted incident response separates signal from noise to slash MTTR and alert fatigue.  
Peter Farago·September 3, 2025

**MTTR: The Emergency Room Metric for SRE**  
MTTR: The Emergency Room Metric for SRE  
Peter Farago·August 26, 2025

**Is Vibe Coding Rewriting Software Development?**  
When AI writes 95% of the code, the engineer's job shifts from production to direction.  
Peter Farago·August 12, 2025

**Your Top Engineer Just Gave Notice**  
From tribal knowledge to antifragility: How to stop a talent exodus from torching your engineering know-how.  
Peter Farago·August 5, 2025

**Why Most Enterprise AI Projects Fail Before They Even Start**  
Stop asking "how to use AI" and start asking which business problems are finally solvable.  
Peter Farago·June 24, 2025

**Beyond the Model Wars: The Real AI Race Begins**  
Applications, not models, will increasingly define the next phase of AI innovation  
Peter Farago·June 17, 2025

**AI's Last Mile Problem**  
Bridging the gap between out-of-the-box LLMs and practical production value.  
Peter Farago·May 20, 2025

**How Corelight Saved 30% of Technical Support Time with RunLLM’s AI Support Engineer**  
RunLLM’s AI Support Engineer helped Corelight’s team save time, respond faster, and maintain quality without adding headcount.  
Peter Farago·May 14, 2025

**How vLLM Uses RunLLM's AI Support Engineer to Deflect 99% of All Technical Questions**  
Deflecting Massive Volume, Freeing Maintainers, Scaling Effortlessly.  
Peter Farago·May 7, 2025

**Beyond AGI: Why Specialization Is the Real AI Breakthrough**  
Beyond AGI: Why Specialization Is the Real AI Breakthrough  
Peter Farago·May 2, 2025

**How DataHub Saved $1MM with RunLLM's AI Support Engineer**  
RunLLM saved DataHub $1 million in engineering cost, increased question capacity 6X, and improved ticket deflection by 90%.  
Peter Farago·April 30, 2025

**Arize AI Transforms Technical Support with RunLLM**  
50% Faster Resolutions, 25% Less Engineering Work, and a 15% Boost in Customer Retention.  
RunLLM Team·March 5, 2025

**The Hard Thing About Building AI Applications**  
How we moved beyond the hype to define the 4 core principles of great AI-native design.  
Vikram Sreekanti·February 28, 2025

**So you want to buy your first AI product**  
So you want to buy your first AI product  
Vikram Sreekanti, Joseph E. Gonzalez·February 13, 2025

**DeepSeek, o3, and AI applications**  
DeepSeek, o3, and AI applications  
Vikram Sreekanti and Joseph E. Gonzalez·February 7, 2025

**AI is yet another platform shift**  
AI is yet another platform shift  
Vikram Sreekanti and Joseph E. Gonzalez·January 30, 2025

**One month of using Devin**  
One month of using Devin  
Vikram Sreekanti and Joseph E. Gonzalez·January 23, 2025

**The end of scaling laws doesn't matter**  
The end of scaling laws doesn't matter  
Vikram Sreekanti and Joseph E. Gonzalez·December 5, 2024

**Your AI strategy is a waste of time**  
Your AI strategy is a waste of time  
Vikram Sreekanti and Joseph E. Gonzalez·November 21, 2024

**A theory of the AI market**  
A theory of the AI market  
Vikram Sreekanti and Joseph E. Gonzalez·October 17, 2024

**In defense of vibes-based evaluations**  
In defense of vibes-based evaluations  
Vikram Sreekanti and Joseph E. Gonzalez·August 1, 2024

**LLMs are becoming commodities**  
LLMs are becoming commodities  
Vikram Sreekanti and Joseph E. Gonzalez·May 30, 2024

**You can't build a moat with AI**  
You can't build a moat with AI  
Vikram Sreekanti and Joseph E. Gonzalez·April 11, 2024

**OpenAI is too cheap to beat**  
OpenAI is too cheap to beat  
Vikram Sreekanti and Joseph E. Gonzalez·October 12, 2023

**RLHF and LLM evaluations**  
RLHF and LLM evaluations  
Joseph E. Gonzalez·September 14, 2023
