Blog | Herald Blog

From the team

Insights on AI agents, software reliability, and the future of production operations.

What It Actually Takes for AI to Map Your Production Systems
An agent cannot diagnose incidents without understanding your system and how it's connected. Here is what that involves.
Chenggang Wu·August 26, 2026·5 min read

System Understanding Is the Whole Game for AI SREs
Momento CTO Daniela Miao on why an AI SRE can only be trusted with alerting decisions if it understands how a system is built.
Peter Farago·August 18, 2026·5 min read

Engineering Knowledge Replicants
Tribal knowledge isn't collective. It's fragmented and fragile, and incident response depends on it. Why documentation and DIY agents still fall short for AI SREs.
Chenggang Wu·July 28, 2026

Winter is Coming for AI Engineering
Every VP of Engineering I talk to right now is spending money on Claude Code or Cursor. Most of them can't tell me what they're getting for it.
Chenggang Wu·July 15, 2026·6 min read

Why You Should Definitely DIY Dev Tools with AI. Sometimes.
Engineering teams are building their own AI dev tools more than ever. After comparing notes with a lot of them, here's where DIY pays off and where it doesn't.
Chenggang Wu·June 30, 2026·5 min read

ClickHouse + Herald: AI DevOps intelligence on fast, cost-efficient telemetry
Herald, an AI DevOps agent in the House Mates program, puts the telemetry you keep in ClickStack to work: predicting incidents, delivering root cause in minutes, and answering questions about your systems. ClickHouse customers can get up to $25,000 in free Herald usage.
Peter Farago·June 25, 2026

No Runbooks, No Problem: Snorkel AI Gets Day-One Results with Herald
93% accuracy on engineering questions. Incidents resolved in minutes. No runbooks. No RCA documents. No Slack history.
Kartik Mathur·June 4, 2026·4 min read

Heralding the Future
We're excited to share that RunLLM is now Herald.
Vikram Sreekanti·May 27, 2026·4 min read

When Your AI-Powered RCA Spews Pages of Useless Text
When your AI-powered RCA tool floods Slack with hallucinated walls of text, the problem isn't the model. It's the missing data engineering underneath it.
Chenggang Wu·May 19, 2026·7 min read

Could Your AI-Generated Code Destroy Your Company?
When everyone can build software, someone still has to keep it running. A reliability leader helped me understand how engineering organizations are facing a new influx of code from all sides.
Chenggang Wu·May 5, 2026

The Code Nobody Read Is Already in Production
Ben Sigelman argues that AI-generated code is a reliability crisis in slow motion, and what it means for how we observe production systems.
Peter Farago·April 29, 2026

The Future of Software is Production
Ship every piece of code you write directly into production.
Vikram Sreekanti·April 21, 2026

Why LLM-Over-Logs Is the Wrong Abstraction.
Dumping logs into an LLM causes high variance and latency. Learn the data engineering approach for AI SRE that prioritizes signal over context.
Vikram Sreekanti·April 14, 2026

I Don’t Care if AI Wrote the Code. You Own It.
SREcon Chair Heinrich Hartmann on why the age of AI-assisted engineering demands a radical return to design rigor.
Peter Farago·April 7, 2026

The SDLC is Dead. Long Live the SDLC.
You cannot review your way out of the new glut of AI code. Winning teams will learn faster from what reaches production.
Vikram Sreekanti·April 1, 2026

The On-Call Problem AI Can Actually Solve
Heinrich Hartmann argues AI’s most valuable role isn’t autonomous remediation. It’s ensuring on-call engineers have the context they need to fix incidents fast.
Peter Farago·March 10, 2026

AI-Created Code Is Putting Us in Debt
The velocity trap is real. Here is the new engineering framework for surviving the age of AI-generated code.
Peter Farago·October 21, 2025

Can AI Spot Outages Faster Than Your Customers?
How AI shortens detection time and prevents trust-eroding surprises
Peter Farago·October 14, 2025

The End of SRE Tribal Knowledge
How AI turns expert intuition into operational infrastructure
Peter Farago·October 7, 2025

Never Let a Good Incident Go to Waste
How AI turns firefighting into continuous learning
Peter Farago·October 1, 2025

Why SREs Need an AI Teammate
AI that clears the path so on-call engineers move faster
Peter Farago·September 25, 2025

Respecting Control by Design
Principles for adding an AI teammate to incident response
Peter Farago·September 17, 2025

The Glass Box AI SRE
Why Transparency Wins in Incident Response
Peter Farago·September 9, 2025

More Needle, Less Haystack: Solving the AI SRE Trust Gap
How AI-assisted incident response separates signal from noise to slash MTTR and alert fatigue.
Peter Farago·September 3, 2025

MTTR: The Emergency Room Metric for SRE
MTTR: The Emergency Room Metric for SRE
Peter Farago·August 26, 2025

Is Vibe Coding Rewriting Software Development?
When AI writes 95% of the code, the engineer's job shifts from production to direction.
Peter Farago·August 12, 2025

Your Top Engineer Just Gave Notice
From tribal knowledge to antifragility: How to stop a talent exodus from torching your engineering know-how.
Peter Farago·August 5, 2025

Why Most Enterprise AI Projects Fail Before They Even Start
Stop asking "how to use AI" and start asking which business problems are finally solvable.
Peter Farago·June 24, 2025

Beyond the Model Wars: The Real AI Race Begins
Applications, not models, will increasingly define the next phase of AI innovation
Peter Farago·June 17, 2025

AI's Last Mile Problem
Bridging the gap between out-of-the-box LLMs and practical production value.
Peter Farago·May 20, 2025

How Corelight Saved 30% of Technical Support Time with RunLLM’s AI Support Engineer
RunLLM’s AI Support Engineer helped Corelight’s team save time, respond faster, and maintain quality without adding headcount.
Peter Farago·May 14, 2025

How vLLM Uses RunLLM's AI Support Engineer to Deflect 99% of All Technical Questions
Deflecting Massive Volume, Freeing Maintainers, Scaling Effortlessly.
Peter Farago·May 7, 2025

Beyond AGI: Why Specialization Is the Real AI Breakthrough
Beyond AGI: Why Specialization Is the Real AI Breakthrough
Peter Farago·May 2, 2025

How DataHub Saved $1MM with RunLLM's AI Support Engineer
RunLLM saved DataHub $1 million in engineering cost, increased question capacity 6X, and improved ticket deflection by 90%.
Peter Farago·April 30, 2025

Arize AI Transforms Technical Support with RunLLM
50% Faster Resolutions, 25% Less Engineering Work, and a 15% Boost in Customer Retention.
RunLLM Team·March 5, 2025

The Hard Thing About Building AI Applications
How we moved beyond the hype to define the 4 core principles of great AI-native design.
Vikram Sreekanti·February 28, 2025

So you want to buy your first AI product
So you want to buy your first AI product
Vikram Sreekanti, Joseph E. Gonzalez·February 13, 2025

DeepSeek, o3, and AI applications
DeepSeek, o3, and AI applications
Vikram Sreekanti and Joseph E. Gonzalez·February 7, 2025

AI is yet another platform shift
AI is yet another platform shift
Vikram Sreekanti and Joseph E. Gonzalez·January 30, 2025

One month of using Devin
One month of using Devin
Vikram Sreekanti and Joseph E. Gonzalez·January 23, 2025

The end of scaling laws doesn't matter
The end of scaling laws doesn't matter
Vikram Sreekanti and Joseph E. Gonzalez·December 5, 2024

Your AI strategy is a waste of time
Your AI strategy is a waste of time
Vikram Sreekanti and Joseph E. Gonzalez·November 21, 2024

A theory of the AI market
A theory of the AI market
Vikram Sreekanti and Joseph E. Gonzalez·October 17, 2024

In defense of vibes-based evaluations
In defense of vibes-based evaluations
Vikram Sreekanti and Joseph E. Gonzalez·August 1, 2024

LLMs are becoming commodities
LLMs are becoming commodities
Vikram Sreekanti and Joseph E. Gonzalez·May 30, 2024

You can't build a moat with AI
You can't build a moat with AI
Vikram Sreekanti and Joseph E. Gonzalez·April 11, 2024

OpenAI is too cheap to beat
OpenAI is too cheap to beat
Vikram Sreekanti and Joseph E. Gonzalez·October 12, 2023

RLHF and LLM evaluations
RLHF and LLM evaluations
Joseph E. Gonzalez·September 14, 2023