You have 1 article left to read this month before you need to register a free LeadDev.com account.
Estimated reading time: 6 minutes
Key takeaways:
- Meta encoded senior engineers’ reasoning into reusable “skills,” not just a generic tool.
- Separating stable “tools” from evolving “skills” is what enabled the platform to scale.
- Diagnosis time dropped from 10 hours to 30 minutes, with humans still approving production changes.
Minor performance issues here and there might not seem like the end of the world, but at scale, they really add up.
“Working at a hyperscale level, small inefficiencies can compound into inefficient use of compute and power,” Tommy Tran, software engineer at Meta, tells LeadDev.
The thing is, finding, diagnosing, and fixing performance issues can be grating and highly manual. Human expertise doesn’t easily scale to the task, but Meta is showing how AI agents can help.
It is using an internally built agentic platform to revamp engineering operations, specifically around identifying and fixing performance regressions.
According to Meta’s engineering blog, its Capacity Efficiency Program is using agents to recover hundreds of megawatts of power and reduce diagnosis time from about 10 hours to around 30 minutes.
Below, we’ll get a behind-the-scenes look at Meta’s use of agents in software infrastructure optimization.
Your inbox, upgraded.
Receive weekly engineering insights to level up your leadership approach.
The initial problem
Before the agent platform was created, Meta engineers spent significant time investigating software infrastructure inefficiencies. Diagnosing issues often relied on particular senior engineers with deep expertise, who couldn’t be everywhere at once.
“Before I started this project, finding and fixing those inefficiencies depended on a handful of senior efficiency engineers doing it reactively and by hand,” says Tran. Just tracing root causes within a massive interconnected system can take hours.
For instance, certain code patterns can waste CPU cycles, creating enormous waste at scale. “Before, finding and fixing those required a performance expert to manually profile services, trace hot paths, identify the anti-pattern, write a targeted fix, and verify it actually saved resources,” explains Tran.
However, that approach could only reach a fraction of the available opportunities, and could take a specialist hours of work to discover, let alone test and remediate, says Tran. There was a clear opportunity to encode specialist knowledge of these patterns and task AI agents with surfacing and suggesting fixes autonomously.
“The core idea I set out to prove was that you could take the judgment those engineers apply, encode it into a platform, and make it scale,” says Tran. “That is what I built.”
Building a unified agent platform
The first design decision Tran made was to combine proactive system optimization and regression remediation within the same architecture. The second was determining when to use ‘tools’ versus ‘skills.’
“Tools are the stable, reusable capabilities for observing and acting on the system,” he says, adding that he built them on Model Context Protocol (MCP) given that it’s an open standard and helps make the tool layer interoperable.
Skills, on the other hand, are where the domain expertise lives. In agentic development, skills encode instructions and capabilities for agents to follow. In Meta’s platform, they capture task-specific judgment from senior engineers and are intended to evolve and be refined over time.
“I build skills by sitting with senior efficiency engineers and codifying how they actually reason: the checks they run, the signals they trust, the order they investigate in,” says Tran. “It is less ‘write down an answer’ and more ‘encode a playbook.’”
Tran adds that keeping MCP tools and agent skills separate was a conscious architectural choice. “Keeping them separate means the plumbing stays stable while expertise evolves independently, and new problems become new skills rather than new systems.”
A key element is the boundary between agent autonomy and human oversight. “The line I drew is clear,” says Tran. “The platform does the investigation and produces a ready-to-review fix, but a human engineer stays in the loop for anything that changes production.”
More like this
Energy savings and reduced engineering time
By applying this approach, Meta has seen tangible business outcomes, including quantifiable energy savings and reduced engineering time.
Now, the platform handles a volume of efficiency work that would traditionally have required a large team of specialists. “A single agent run covers ground that used to take an engineer days,” says Tran. Also, it works continuously.
“The outcome I care most about is that the system turned what was a reactive, one-case-at-a-time process into continuous, fleet-wide coverage operating around the clock,” says Tran. Those results validated the approach, and the same kind of gains are available to any organization managing infrastructure at scale.
Usability is also enhanced. A domain expert can describe an inefficiency pattern once as a natural-language prompt, and the platform turns that into an agent that scans relevant code, identifies every instance, and produces ready-to-review code changes. That, in turn, frees up engineers for other meaningful work.
“It fundamentally changed what engineers spend their time on,” he adds. “Work that used to be hours of manual investigation is now reviewing a proposed fix, which frees those engineers for harder, novel problems that actually require human creativity.”
Tips for other agentic system builders
Meta’s story comes at a moment when many engineering teams are investigating how AI agent tools can benefit their workflows. Some are in the process of building their own agentic systems.
However, teams no longer have to go into this blind: Meta’s case study sheds some light on helpful takeaways for engineering leadership.
Ground the project in a real use case
“First, start with a real, painful, well-understood workflow, not a demo,” advises Tran. In his case, it was proven operational enhancements, but the same philosophy could be applied in many other areas.
Separate tools from skills
Second, consider which capabilities should remain stable tools versus evolving skills. “That single architectural decision is what lets an agent platform generalize instead of ossifying,” says Tran.
Don’t replace engineers, enhance them
Meta’s agent platform didn’t remove engineers – it enhanced them. In Meta’s framework, engineers remain in control, and the goal is to remove the hours of manual investigation, not the engineer’s judgment. So, consider how new agent flows can enhance existing talent, not replace it.
Encourage internal use
Getting a project like this off the ground, especially one that changes engineering workflows, will take convincing. “What helped most was showing measurable wins on real cases rather than arguing in the abstract,” says Tran. The fact that the platform produces low-risk suggestions helped make the case, too.
Get senior engineers involved
Tran also got senior engineers involved early on in the process, which aided trust and benefited adoption. “Once the people whose judgment was being encoded saw the system reflect their own reasoning, they trusted it. That trust is what drives adoption. I think that principle holds anywhere you are introducing AI into an engineering workflow, not just at Meta.”
Keep humans in the loop
Meta’s system does the hard work around discovery and makes the output easy to review. It doesn’t risk making production deployments without human judgment.
Measure the impact
Finally, frame the results in business terms. “Measure impact in terms leadership already cares about: cost, energy, engineering time,” recommends Tran.

New York • September 8 & 9, 2027
Loved LDX3 New York? Pre-sale tickets for 2027 are now available.
Out of a few people’s heads
Many large organizations have deep expert knowledge, but that knowledge doesn’t always scale across the organization. “Every field has senior experts whose judgment is trapped in their heads,” says Tran.
Perhaps the most relatable outcome of this approach is the ability to share engineering knowledge through agent skills. For Tran, this offers a repeatable method for unlocking and scaling expertise across an organization.
Since the skills encode expert knowledge, such a framework could also prevent institutional knowledge decay, a common problem in software engineering.
“It broke deep efficiency knowledge out of a few people’s heads and embedded it in the platform where the whole organization benefits,” says Tran. “That shift is what I think other engineering organizations will recognize in their own operations.”