An Agentic Mixture of Experts for DevOps with Sunil Mallya -

The TWIML AI Podcast (formerly This Week in...

An Agentic Mixture of Experts for DevOps with Sunil Mallya

Listen now

Description

Today we're joined by Sunil Mallya, CTO and co-founder of Flip AI. We discuss Flip’s incident debugging system for DevOps, which was built using a custom mixture of experts (MoE) large language model (LLM) trained on a novel "CoMELT" observability dataset which combines traditional MELT data—metrics, events, logs, and traces—with code to efficiently identify root failure causes in complex software systems. We discuss the challenges of integrating time-series data with LLMs and their multi-decoder architecture designed for this purpose. Sunil describes their system's agent-based design, focusing on clear roles and boundaries to ensure reliability. We examine their "chaos gym," a reinforcement learning environment used for testing and improving the system's robustness. Finally, we discuss the practical considerations of deploying such a system at scale in diverse environments and much more. The complete show notes for this episode can be found at https://twimlai.com/go/708.

More Episodes

See all »

Automated Reasoning to Prevent LLM Hallucination with Byron Cook

Today, we're joined by Byron Cook, VP and distinguished scientist in the Automated Reasoning Group at AWS to dig into the underlying technology behind the newly announced Automated Reasoning Checks feature of Amazon Bedrock Guardrails. Automated Reasoning Checks uses mathematical proofs to help...

Published 12/09/24

AI at the Edge: Qualcomm AI Research at NeurIPS 2024 with Arash Behboodi

Today, we're joined by Arash Behboodi, director of engineering at Qualcomm AI Research to discuss the papers and workshops Qualcomm will be presenting at this year’s NeurIPS conference. We dig into the challenges and opportunities presented by differentiable simulation in wireless systems, the...

Published 12/03/24

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

Published 12/03/24