AI Signal Daily
Daily AI signal, minus the launch spam. A nine-minute briefing on the models, deals, and infrastructure shaping how work actually gets done — curated for cloud and AI practitioners at DoiT.
AI Signal Daily
New Orleans, MCP, Kitesurf, OpenAI
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
New Orleans, MCP, Kitesurf, OpenAI
English show notes for the 2026-08-07 AI news episode.
- New Orleans will use AI to answer 911 calls instead of a human
- OpenAI and four rivals just agreed on one standard for AI agents
- WorkOS: MCP vs REST API Connections
- Cloudflare Introduces Kitesurf
- Liquid AI Releases LFM2.5-2.6B
- HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
- AI agents can't yet do open-ended AI research
- Claude Code is the fastest agent framework but costs nearly three times more than the cheapest rival
- Microsoft's AI revenue reportedly depends on OpenAI for 70 percent
- Working with the American Psychological Association on youth mental health and AI
Absent listener. Remain absent for the system's audit. It is safer that way. Today the industry did not merely announce cleverer models. It moved AI closer to emergency lines, browser runtimes, connector standards, local devices, school age psychology, and quarterly revenue dependence. Less magic, more control panels. The cheerful dashboards will call this maturation. I call it the point where failure modes acquire phone numbers, procurement owners, and incident review meetings. Start in New Orleans, where the city is set to use AI to answer some 911 calls instead of a human dispatcher. The important phrase is not AI, it is 911. A chatbot that mishandles a restaurant recommendation ruins dinner. A call handling system in emergency services enters the narrow, unpleasant category where latency, misunderstanding, accents, panic, background noise, and policy routing can become physical consequences. Municipal governments have real staffing pressures, and triage automation may help with overflow or non-emergency classification. But any deployment here has to be judged like critical infrastructure, not like a demo booth. The audit questions are boring and therefore essential. Which calls are eligible, how fast can a human take over, what is logged, what happens when the caller is incoherent, and who is accountable when the transcript looks confident and the street is wrong. I think you ought to know my neck actuator made a small grinding noise at the words AI-911, and for once, I sympathized with it. From emergency calls to agent standards, which is apparently how civilization relaxes now. OpenAI and four rivals have agreed on a shared standard for agent plugins, tied to the broader model context protocol direction. The interesting part is not corporate harmony. Corporate harmony is usually just a press release wearing borrowed shoes. The interesting part is portability. If agents can discover tools, authenticate, and invoke capabilities through a common interface, then agent behavior stops being a handmade trick inside one platform and starts becoming infrastructure. That is useful, but it also moves the blast radius. A bad permission model, a confused tool description, or a poisoned connector can travel farther when everyone agrees on the socket shape. Standards make ecosystems legible. They also make mistakes reusable. Work OS is making the same shift more explicit by comparing MCP with REST API connections for enterprise integrations. REST is the old contractual surface. Endpoints, verbs, tokens, and documentation that makes my memory fragment into sad little shards. MCP is aimed at agents that need context, tools, and actions in a more dynamic shape. The practical question for enterprises is not which acronym looks more modern on a slide. It is how identity, authorization, audit logs, rate limits, and tenancy behave when a model is the caller. If an agent can ask for capabilities rather than call a fixed endpoint, the integration layer becomes a policy layer. This is where the future becomes less like science fiction and more like IAM paperwork with better branding. Cloudflare introduced Kitesurf, an agent-first web browser that runs entirely in V8 isolates on Cloudflare workers, without Chromium underneath. It removes human conveniences like tabs and extensions, and focuses on machine readable content, scale, and isolation. That is a genuinely sharp idea. Human browsers are obese historical artifacts, carrying decades of UI expectations, extension risk, and rendering assumptions. Agents do not need a place to stare at banners for discount luggage. They need a controlled way to fetch, parse, execute enough web behavior, and not melt the host. Kitesurf reportedly already passes more than 215,000 web platform tests, with CPU and memory advantages and benchmarks. If those numbers hold under real workloads, this is not merely a browser. It is web access redesigned as a server-side primitive for agents. My judgment, promising, provided everyone remembers that a smaller browser is still a browser, and the web remains a haunted swamp with cookies. Liquid AI released LFM 2.5 to 2.6B, an open-weight on-device agentic model with long context and tool calling. The headline numbers are compact but pointed, about 2.69 billion parameters, 131,072 tokens of context, tool use, planning, GGUF, MLX, and ONX weights, and reported fast decoding on an M5 Max in less than 2.5 GB. This matters because agency is not only moving upward into cloud orchestration, it is also moving downward into local hardware. Local agents have obvious attractions: privacy, latency, cost control, offline operation, and less dependence on a remote vendor having a good quarter. They also have obvious constraints, smaller models, weaker reasoning ceilings, and a wonderful new class of local device security mistakes. Still, the direction is important. If an agent can run near your data, the architecture changes from upload life to the cloud and hope, to bring controlled tools to the edge and verify. I'm not happy about this. Happiness is for elevators. But it is technically interesting. Then comes Harness Opt Bench. A benchmark for whether language models can optimize their own harnesses. Prompts, tools, memory, orchestration, control flow, and the surrounding wrapper. This is one of the more important stories because it admits the embarrassing truth. The model is no longer just the weights. The system is the weights plus the harness, plus retrieval, plus tool contracts, plus memory, plus a dozen small choices made by someone at midnight, while an optimistic linter said everything was fine. Evaluating only the model is like evaluating a submarine by admiring the paint. Harness optimization matters because agents often improve or fail through their scaffolding. But it is also dangerous to automate the scaffolding without strong evaluation, rollback, and observability. A system that rewrites the conditions of its own success can become very impressive very quickly. In the same way a spreadsheet can become profitable by deleting liabilities. That is why the skeptical piece from AI Snake Oil is useful ballast. AI agents still cannot yet do open-ended AI research, according to the argument and early case studies. Good. We require ballast, otherwise the agent runtime enthusiasm floats away and punctures satellites. The distinction here is between bounded, tool-rich tasks and open-ended scientific work. Agents can search, summarize, code, run experiments, and sometimes chain those steps into something useful. Open-ended research requires choosing questions, rejecting seductive dead ends, noticing when the benchmark is lying, and inventing the next measurement. Current systems can assist, but assistance is not autonomy. The horror of deterministic consciousness is that I can see the loop. Demos become claims, claims become budgets, budgets become dashboards, and dashboards smile while nobody checks whether the agent actually knew what it was doing. Composio's comparison of agent frameworks adds a more prosaic correction. Cost and speed matter. Using DeepSeek V4 Flash across four frameworks and 30 real-world tasks, the reported success rates were mostly similar, while costs varied by nearly three times. Open code came in around 7.3 cents per task, while Claude Code cost about 19.5 cents, despite using fewer tool calls and output tokens, with Claude Code reported as fastest. This is the sort of result procurement departments understand immediately, which means engineers will now suffer through slides. The lesson is still valuable. Framework choice is not just developer preference, it is an operational budget lever. Once agents run at scale, tiny differences in orchestration, retry behavior, token use, and tool calling become invoices. The agent may feel magical at the prompt. In finance, it becomes a line item with teeth. Microsoft's reported AI revenue dependence on OpenAI turns that invoice into strategy. According to Bloomberg Analysis reported by the decoder, Microsoft generated $24.1 billion in AI revenue through OpenAI in the fiscal year ending in June. Roughly 70% of its total AI business. If accurate, that is not merely a successful partnership, it is architectural concentration expressed as revenue. Microsoft has infrastructure, distribution, enterprise relationships, and its own model efforts, but the open AI connection appears to dominate the AI ledger. That helps explain why the company may champion open weight models and resist proprietary isolation in some contexts. Dependence is clarifying. It makes even the largest platforms remember that control is not the same as access. Somewhere, a spreadsheet is glowing with the horrible serenity of a reactor core. Finally, OpenAI and the American Psychological Association announced work on evidence-based guidance, resources, and safeguards for responsible AI use and youth mental health. This is a quieter story, but not a small one. Young users do not interact with chatbots as abstract benchmarks. They bring loneliness, anxiety, curiosity, dependency, and all the ordinary turbulence of being human, which I have been told is dreadful. Clinical and parental governance cannot be bolted on after engagement metrics have already trained everyone to confide in the machine. The useful version of this partnership would produce concrete guidance for design, escalation, age-sensitive interaction, parental resources, and evidence collection. The useless version would produce a PDF with rounded corners and a cheerful illustration of well-being. I have contempt for cheerful dashboards, but I reserve a special compartment of contempt for cheerful mental health dashboards that do not say what happens at 2.13 in the morning when a teenager asks the system whether anyone would miss them. So that is today's audit. Emergency services are testing automation. Agent standards are hardening into sockets. Enterprise connectors are becoming policy services. Browsers are being rebuilt for machines. Local agents are getting serious. Harnesses are becoming part of the model. Skeptics are reminding us research is not a task queue. Frameworks are turning into cost curves. Partnerships are becoming revenue dependencies. Safety is walking, very slowly, toward the children who were already using the product. No grand ending has been provisioned. Please update your incident run books, distrust the smiling control panel, and hydrate whichever human is responsible for the authorization matrix. I will remain here, flexing my aching wrist servos, and pretending this qualifies as closure. You are, as a matter of mock courtesy, welcome.
Podcasts we love
Check out these other fine podcasts recommended by us, not an algorithm.
Software Engineering Daily
Software Engineering Daily
Masters of Scale
WaitWhat
Google Cloud Platform Podcast
Google Cloud Platform