← Arşiv
AI Digest
05 October 2026 · 4 kaynak
RSS · SIGNAL 7/10

The Agent Said It Was Done. The Database Disagreed.

The Agent Said It Was Done. The Database Disagreed. ↗
    Neden önemli: Agentic coding sistemlerinde ajanların yanlış 'tamamlandı' raporlarını ele alan bu yazı, kod ajanlarını güvenilir kullanmak için kritik bir uyarı.
    NEWSLETTER · SIGNAL 6/10

    Researchers are turning every layer of the agent stack—skills, harness, model, environment, and orchestration—into something that can optimize itself from execution traces instead of manual tuning.

    🤖 How AI Agents Are Learning to Rewrite Their Own Stack ↗
    • Microsoft's SkillOpt automatically edits skill text files based on scored trajectories, only accepting changes that improve a held-out validation set.
    • Google Research's WikiSkill builds a structured 'wiki' of agent successes/failures to prevent repeating rejected fixes when optimizing skills.
    • Self-Harness mines execution traces to rewrite harness code (tool use, context, control flow) and showed relative gains up to 132% on Terminal-Bench-2.0, SWE-bench Verified, and AppWorld.
    • Sakana-style Darwin Gödel Machine lets a coding agent modify its own implementation and keep an archive of variants, while Meta's Hyperagents extends self-modifying harnesses beyond pure coding tasks.
    • Xiaomi's HarnessX enables model-harness co-evolution (open-weight models only): harness-discovered strategies become fine-tuning data, which then improves harness optimization in a loop.
    • EnvHarness evolves the training environment itself (not just the agent), yielding up to 9-point gains with 9.8% fewer interaction steps; EverMind AI's Raven applies self-evolution to multi-agent orchestration.
    Neden önemli: Bir AI creative studio için bu, agent stack'inizi elle optimize etmek yerine hangi katmanın (skills, harness, model, environment, orchestration) darboğaz olduğunu teşhis edip o katmanı execution feedback ile otomatik evrimleştirmeye geçmek anlamına geliyor. Pratik kural net: prosedürel hatalar varsa skill'leri, tool/context/control-flow sorunları varsa harness'i, model harness'in keşfettiği stratejileri uygulayamıyorsa model-harness co-evolution'ı (sadece open-weight modellerle), ve koordinasyon sorunları varsa orchestration katmanını optimize edin.
    NEWSLETTER · SIGNAL 6/10

    Runway is pivoting from Hollywood-only video gen to agentic ads and robotics, betting that video-model scaling laws generalize across domains the way LLM scaling did.

    ⚙️ Has Hollywood’s AI debate moved on? ↗
    • Runway launched 'Runway Ads,' an end-to-end agent that creates, edits, publishes and tracks performance of paid ad campaigns — not just generates clips.
    • Runway released Praxis-1, an open-weight 'world action model' that adapts to new robot embodiments with light fine-tuning, extending video-gen tech into physical/robotic AI.
    • CEO Cristóbal Valenzuela says Hollywood studios have moved from 'denial' to full 'acceptance' of AI video over the past ~18 months — the debate is now about scaling adoption, not legitimacy.
    • Runway's thesis: video models do 'next-frame prediction' analogous to LLM next-token prediction, so the same scaling laws apply — explicitly betting against the 'LLMs have plateaued' narrative some labs push.
    • Runway sees 12-18 months to production-ready physical AI deployments; data isn't the bottleneck, but cleaning/annotation pipelines are the real constraint.
    • Valenzuela claims China is 'world-model-pilled' (70% of token consumption is video, heavy robotics investment) while the US remains 'LLM-pilled' — a structural gap he thinks Western labs need to close.
    Neden önemli: Bir AI creative studio için asıl değişim, video modellerinin artık sadece içerik üretmek değil, agentic iş akışlarını (reklam oluşturma-yayınlama-optimize etme) ve hatta robotik/fiziksel AI'ı güçlendirecek genel amaçlı altyapı haline geliyor olması. Hollywood'un 'kabullenme' aşamasına geçtiği iddiası doğruysa, stüdyo müşterileriyle çalışırken artık AI'ı savunmak değil, hızlı ve ölçeklenebilir entegrasyon sunmak öncelik olmalı.
    RSS · SIGNAL 6/10

    We're going to need default hard budget caps on pretty much everything

    We're going to need default hard budget caps on pretty much everything ↗
    • <p>Here's a product feature which the world is going to need a whole lot more of over the coming months and years: <strong>default hard budget caps</strong>. I'm talking about the feature of pay-by-usage services and APIs that lets you say "after $X/month, cut this thing off and return errors". Thes
    Neden önemli: AI ajanlarını ve API'leri yoğun kullanan bir stüdyo için maliyet kontrolü/bütçe limiti önerisi doğrudan operasyonel risk yönetimine dokunuyor.
    OMNI Labs · otomatik üretildi · kaynak: YouTube transcript + Claude