558 episodios
- What happens to an MMORPG when half the players never log off? Issac Lee, who runs AI and blockchain experiments at NEXUS, has been testing that idea. His team is filling game worlds with agent players instead of scripted NPCs, so a new title feels alive from day one.
We get into why agent-generated content could be the next step after studio-made and user-generated games. We also cover how agent populations might solve the chicken-and-egg problem every multiplayer launch faces, and what it means when your fellow players are relentless and never need sleep. Isaac walks through the NEXUS lab: a text-to-3D world builder for Roblox, an AI pipeline that turns ONE store novels into short web dramas, a loop-based tool for natural pixel animation, and Da Vinci, the internal multi-agent framework the team uses to ship.
Then the conversation turns to infrastructure. We talk about why NEXUS went from 200 hand-defined agent roles to a much thinner harness, whether custom harnesses still matter now that Claude and Codex exist, and why the Kimi release pushed the company to start building its own small data center for inference. We close on context: how much of your life you should hand to an AI, what personalized meeting notes could look like, and why shopping agents may force every ecommerce site to rethink who it's designing for.
Demetrios Brinkmann: https://www.linkedin.com/in/dpbrinkmIssac Lee: https://www.linkedin.com/in/issacley/?locale=ko
Timestamps:[0:00] Cold open[0:57] AI in gaming at NEXUS[2:19] Text to 3D and the last 10 percent[4:38] AI companions and coaches[6:35] Agent players in MMORPGs[8:27] From UGC to agent-generated content[11:07] Agents that never sleep[12:58] Taming fuzzy prompts[19:24] Inside Da Vinci[23:10] Undercover QA agents[27:14] Do harnesses still matter[31:52] Self-hosting open models[41:33] Lab experiments and Bob[50:42] How much context is too much[54:22] Marketing and shopping agents - Picture a startup with ten engineers hammering away on AI and one person lying awake over a $20,000 bill that hasn't arrived yet. That's the scenario we put to Jason Ward, who handles FinOps for AI at C.H. Robinson and recently joined the FinOps Foundation's AI working group.
Jason's first answer is not glamorous: tag every AI resource so you know who owns it. The rest of the episode is what that makes possible. He walks through how C.H. Robinson runs AI across order entry, quoting, booking, and tracking, and why the team routes Anthropic models through Vertex AI to keep data locked down.
Then come the metrics. Jason uses AI to dig through his own observability platform for signals he didn't know were there. One of them is how chatty a model is. That signal turned a prompt bloat alert into a bug in the code that kept retrying and burning tokens. Its opposite, context starvation, burns tokens too: a model with too little context keeps failing and trying again.
We also cover why agentic and conversational workloads need separate baselines. Jason explains why cost per order is the easy win, and why most of the real work doesn't fit into neat discrete tasks. That's where his experimental cost per thought metric comes in, with reasoning ratio and cache hit rate alongside it. His advice is simple: your AI is the best tool you have for understanding your AI.
Alex Salkever: https://www.linkedin.com/in/alexsalkever
Jason Ward: https://www.linkedin.com/in/jward2
Timestamps:[0:00] Cold open[0:33] Meet Jason Ward from C.H. Robinson[1:18] AI use cases in logistics[1:52] Azure OpenAI and Vertex AI[2:19] Why keeping data in-house matters[2:44] How developers use AI day to day[3:40] Using AI to hack AI observability[4:21] Measuring how chatty a model is[5:45] Advice for a startup afraid of its AI bill[6:33] Step one is tag every AI resource[7:05] The Copilot billing blind spot[7:44] Why spend is only half the story[8:12] Break down cost by app and by model[9:09] The prompt bloat that exposed a bug[10:03] Context starvation[11:28] What goes into the AI spend report[12:21] Discrete tasks are low-hanging fruit[12:58] Agentic vs conversational workloads[14:12] Know the workflow before you report on it[15:02] The CTO's end goal[16:08] Cost per thought[17:59] Reasoning ratio and cache hit rate[19:06] Dynamic model routing[20:04] Use your AI to improve your AI - Caveman prompting has one rule: why use many words when few do the trick? It saves tokens on the way in and on the way out. Push it too far, though, and the output falls apart. So how far is too far? Nobody has benchmarked it yet, and that question opens our conversation with James Barney, Head of Forward Labs at MetLife.
James spends his days connecting new AI capabilities to old business problems across dozens of regulatory regimes, and he still finds time to push code. He explains how the FinOps Foundation's AI working group took on the most basic question: which model for which workload, and why the answer always comes down to cost, speed, and accuracy. We get into Anthropic's launch pricing for Fable, why a million tokens is easy to price and hard to explain, and why every stakeholder eventually tells you what they really care about once you name the wrong North Star.
From there it gets practical. Start with the smartest model, then step down and add harness until quality holds. Treat exploration tokens like local builds and production tokens like pipelines. Govern agents the way you govern people, with proactive blocks, reactive checks, and policy as code an agent can actually read.
We close on a bigger shift. When a chat window can pull from every dashboard at once, do we still need dashboards? James thinks mostly not, with one catch he calls the latent shopper problem: some insights only come from browsing data you did not know to ask about.
Demetrios Brinkmann: https://www.linkedin.com/in/dpbrinkm
James Barney: https://www.linkedin.com/in/james-barney
Timestamps:
[00:00] Cold open
[01:03] Meet James from MetLife
[01:11] What AI enablement means at a global insurer
[03:27] Inside the FinOps Foundation AI working group
[05:03] The caveman skill challenge
[06:37] Why we need a CaveBench
[07:43] Anthropic's Fable launch pricing
[09:34] Planning for surprise model releases
[11:45] Measuring AI value beyond cost
[12:47] Finding your unit metric
[14:15] Tying token spend to business outcomes
[16:35] Explore first, then optimize
[18:12] AI that tunes itself
[20:13] Do R&D tokens count
[23:24] When the experiment becomes the product
[25:38] Personal agents and daily briefings
[27:39] Governing agents that run just because they can
[30:37] The layers of AI governance
[31:56] Proactive and reactive guardrails
[32:59] Policy as code for agents
[34:41] Stop the click ops
[35:31] Is the modern UI obsolete
[36:36] MCP apps and chat as the new browser
[38:07] How AI gathers data differently than humans
[40:37] The latent shopper problem
[42:12] Staying close to your data - What happens after you’ve built your first MCP server and actually have to make it work in the real world?
In this episode, James Ward dives into the more advanced side of MCP: observability, evals, tool design, code mode, authentication, and the challenges that appear once agents start using your server at scale.
We also get into how AWS thinks about MCP across roughly 16,000 APIs, why inefficient tool design often gets blamed on MCP itself, and whether the future could involve more constrained, human-reviewable alternatives to full code mode.
Along the way, James shares a great example of an AWS documentation change that accidentally triggered prompt injection warnings from agents, showing just how complicated testing across different models and harnesses is becoming.
If you’re already building with MCP and want to understand what comes after the “hello world” stage, this one goes deep.
Timestamps:
[00:00] “This Code Is Gobbledygook”
[00:44] What Happens After You Build an MCP Server?
[02:11] The MCP Patterns You Actually Need in Production
[03:48] Why Your Agent Is Making Too Many Tool Calls
[05:29] How to Make MCP Use Fewer Tokens
[08:44] Is MCP Actually Inefficient?
[10:11] “We Built a Lot of Pretty Crappy MCP Servers”
[11:25] The MCP 2.0 Migration Problem
[14:02] Why AWS Is Rethinking Code Mode
[15:35] The Problem With Letting Agents Write Python
[19:31] How Do You Measure Agent Experience?
[22:11] Finding Out Why an Agent Failed
[23:42] Hundreds of Evals for Four Cents
[24:36] The Evaluation Matrix Gets Massive
[26:28] AWS Accidentally Triggered a Prompt Injection Warning
[30:16] Should MCP Servers Expose Only Five Tools?
[31:02] AWS Has 16,000 APIs. Now What?
[34:09] One Super-Agent or Thousands of Specialized Agents?
[36:55] The MCP Authentication Problem
[40:15] What’s Coming at AgentCon - Tool descriptions tell an agent what a tool does. They don't tell it how to use five tools together, in the right order, following your conventions. That gap is where this conversation lives.
Filmed at AGNTCon + MCPCon in Tokyo with Ola Hungerford, Principal Engineer for AI Enablement at Nordstrom and a maintainer of the Model Context Protocol, who spent the last several months turning a pattern everyone was quietly reinventing into an actual MCP extension.
Ola walks through what skills over MCP really means: the server stops being a pile of tools and becomes a distribution channel, handing the agent the instructions, workflows and knowledge it needs only at the moment it needs them. She explains why server instructions weren't enough, how progressive discovery keeps context from exploding, and why the same mechanism works for memory and preferences even when no tools are involved.Then it gets into the harder parts. What belongs in the MCP spec versus the agent skills spec. Why passing custom front matter through opens a rug pull and prompt injection surface nobody wanted. Where skills start to look like sub-agents, and why there's still no standard way to declare which servers a skill depends on. And the honest problem underneath all of it: how do you standardize something while everyone is still finding out what it's actually for, without breaking a hundred things the next time you change your mind?
Timestamps:[0:00] Intro[0:21] AI enablement at Nordstrom[0:33] What skills over MCP actually is[1:29] The MCP server as a distribution channel[2:01] Server instructions versus skills[3:11] Distributing knowledge and memory[4:23] Progressive discovery explained[5:23] Where the idea came from[6:59] From draft to official extension[8:19] What early adopters changed[8:51] Front matter and custom metadata[9:57] Rug pulls and prompt injection risk[10:57] Will any of this get standardized[12:10] Marrying two very different specs[13:03] Skills as personas and sub-agents[13:53] The missing dependency standard[15:15] How the extension actually works[16:19] What harnesses still need to support[16:59] Consent and skill integrity[17:56] Where skills over MCP goes next[18:43] Why cramming 200 tools fails[19:41] Standardizing before you know the answer[22:18] Is git the wrong tool for agents[23:19] Picking tools for the actual persona[23:51] Trying to be less productive[25:48] The anxiety of idle agents[27:15] Why she keeps a robot on her desk[28:22] If the agent feels the friction, does it matter[29:41] Efficiency, waste, and caring enough[30:52] Letting an agent debug for you[32:03] Choosing your rabbit hole
Más podcasts de Tecnología
Podcasts a la moda de Tecnología
Acerca de Agentic Conversations (formally mlops.community)
Relaxed conversations and technical deep dives around AI Agents. This Show is brought to you by the Agentic AI Foundation where the leading agentic open-source projects like MCP, Agents.md, and Goose live. See more at aaif.io
Sitio web del podcastEscucha Agentic Conversations (formally mlops.community), Mundo Futuro y muchos más podcasts de todo el mundo con la aplicación de radio.net

Descarga la app gratuita: radio.net
- Añadir radios y podcasts a favoritos
- Transmisión por Wi-Fi y Bluetooth
- Carplay & Android Auto compatible
- Muchas otras funciones de la app
Descarga la app gratuita: radio.net
- Añadir radios y podcasts a favoritos
- Transmisión por Wi-Fi y Bluetooth
- Carplay & Android Auto compatible
- Muchas otras funciones de la app


Agentic Conversations (formally mlops.community)
Escanea el código,
Descarga la app,
Escucha.
Descarga la app,
Escucha.
Agentic Conversations (formally mlops.community): Podcasts del grupo



























