← Agentic Conversations (formally mlops.community)

The Caveman Prompting Challenge

Agentic Conversations (formally mlops.community)2026年10月1日44分

The Caveman Prompting Challenge

Agentic Conversations (formally mlops.community)

0:0044:07
このエピソードの日本語要約を準備中です。
番組の概要欄(原文)

<p>Caveman prompting has one rule: why use many words when few do the trick? It saves tokens on the way in and on the way out. Push it too far, though, and the output falls apart. So how far is too far? Nobody has benchmarked it yet, and that question opens our conversation with <em>James Barney</em>, Head of Forward Labs at <a href="https://www.metlife.com/" target="_blank" rel="ugc noopener noreferrer">MetLife</a>.</p><p><br></p><p>James spends his days connecting new AI capabilities to old business problems across dozens of regulatory regimes, and he still finds time to push code. He explains how the FinOps Foundation&#39;s AI working group took on the most basic question: which model for which workload, and why the answer always comes down to cost, speed, and accuracy. We get into Anthropic&#39;s launch pricing for Fable, why a million tokens is easy to price and hard to explain, and why every stakeholder eventually tells you what they really care about once you name the wrong North Star.</p><p><br></p><p>From there it gets practical. Start with the smartest model, then step down and add harness until quality holds. Treat exploration tokens like local builds and production tokens like pipelines. Govern agents the way you govern people, with proactive blocks, reactive checks, and policy as code an agent can actually read.</p><p><br></p><p>We close on a bigger shift. When a chat window can pull from every dashboard at once, do we still need dashboards? James thinks mostly not, with one catch he calls the latent shopper problem: some insights only come from browsing data you did not know to ask about.</p><p><br></p><p>Demetrios Brinkmann: <a href="https://www.linkedin.com/in/dpbrinkm" target="_blank" rel="noopener noreferer">https://www.linkedin.com/in/dpbrinkm</a></p><p>James Barney: <a href="https://www.linkedin.com/in/james-barney" target="_blank" rel="noopener noreferer">https://www.linkedin.com/in/james-barney</a></p><p><br></p><p>Timestamps:</p><p>[00:00] Cold open</p><p>[01:03] Meet James from MetLife</p><p>[01:11] What AI enablement means at a global insurer</p><p>[03:27] Inside the FinOps Foundation AI working group</p><p>[05:03] The caveman skill challenge</p><p>[06:37] Why we need a CaveBench</p><p>[07:43] Anthropic&#39;s Fable launch pricing</p><p>[09:34] Planning for surprise model releases</p><p>[11:45] Measuring AI value beyond cost</p><p>[12:47] Finding your unit metric</p><p>[14:15] Tying token spend to business outcomes</p><p>[16:35] Explore first, then optimize</p><p>[18:12] AI that tunes itself</p><p>[20:13] Do R&amp;D tokens count</p><p>[23:24] When the experiment becomes the product</p><p>[25:38] Personal agents and daily briefings</p><p>[27:39] Governing agents that run just because they can</p><p>[30:37] The layers of AI governance</p><p>[31:56] Proactive and reactive guardrails</p><p>[32:59] Policy as code for agents</p><p>[34:41] Stop the click ops</p><p>[35:31] Is the modern UI obsolete</p><p>[36:36] MCP apps and chat as the new browser</p><p>[38:07] How AI gathers data differently than humans</p><p>[40:37] The latent shopper problem</p><p>[42:12] Staying close to your data</p>

X でシェアSpotify で聴くApple Podcasts で聴く