← Arşiv
AI Digest
07 September 2026 · 4 kaynak
⚠ Kaynak uyarisi
  • HuggingFace Blog (rss) · 5 gundur sessiz
Feed bozulmus olabilir: site yeniden yapilmis, URL 404 donuyor, ya da gonderen adresi degismis olabilir. Kontrol et.
NEWSLETTER · SIGNAL 8/10

Shopify fine-tuned a 0.8B Qwen3.5 model to beat GPT-5.6 Sol xhigh on buyer-profile generation, cutting inference cost from $27M to $1M/year on its GraphQL agent while boosting throughput 36x.

⚙️ What Developers Can Learn From Shopify’s Self-Improving AI Pipeline ↗
  • Shopify's fine-tuned Qwen3.5 0.8B model outperformed GPT-5.6 Sol xhigh on buyer-profile generation, cutting the system prompt from 9,100 to 1,100 tokens and raising throughput from 2M to 72M outputs/day.
  • The Sidekick GraphQL agent flywheel: production conversations are scored by LLM judges, frontier models critique low-scoring 'hard negative' trajectories, repairs are replayed and validated, then fed back as training data via SFT + GRPO.
  • Running Sidekick's traffic (up to 2,000 req/min) on frontier models would cost ~$27M/year vs ~$1M for the specialized model, with 19% faster time-to-first-token, 38% lower end-to-end latency, and 14% fewer GPUs needed.
  • Evaluation is the hard part: Shopify calibrates LLM judges against human-labeled data using Cohen's kappa, and requires 4-judge agreement before accepting automated labels (e.g., for correct-refusal cases missing from production data).
  • Shopify's explicit guidance: don't fine-tune during the prototype/exploration phase — specialization only pays off once a workload is high-volume, bounded, measurably evaluable, and backed by proprietary data.
  • The durable asset isn't the model itself but the loop (data, evals, training recipes, deployment infra) — each new frontier model generation becomes a better 'teacher' for training the next specialist.
Neden önemli: Bir AI creative studio icin bu, prototip asamasinda frontier modellere (orn. GPT-5.6, Claude) baglı kalmanin dogru strateji oldugunu, ama iş belirli ve yuksek hacimli bir gorev haline geldiginde (orn. tekrarlayan icerik uretimi, siniflandirma) kucuk ozel bir modele gecmenin maliyet ve hiz acisindan buyuk kazanc sagladigini gosteriyor. Asil yatirim yapilmasi gereken sey tek bir model degil, uretim hatalarini egitim verisine donusturen degerlendirme ve feedback altyapisi (LLM judge, hard-negative mining, SFT+RL dongusu).
RSS · SIGNAL 8/10

Introducing GPT-6 Astra for developers

Introducing GPT-6 Astra for developers ↗
  • <p><strong><a href="https://www.youtube.com/watch?v=bOC3DisEOfg">Introducing GPT-6 Astra for developers</a></strong></p> Blink and you'll miss it, but there's a familiar creature at <a href="https://www.youtube.com/watch?v=bOC3DisEOfg&amp;t=119">1m59s</a>:</p> <blockquote> <p>Across the board, Astra
Neden önemli: Yeni model erişimi; video/görsel üretim iş akışlarına entegre edilebilecek güçlü bir yapay zeka aracı.
RSS · SIGNAL 8/10

Using Blender with coding agents on macOS

Using Blender with coding agents on macOS ↗
  • <p><strong>TIL:</strong> <a href="https://til.simonwillison.net/llms/blender-coding-agents-macos">Using Blender with coding agents on macOS</a></p> <p>I've been having fun with Blender in ChatGPT Codex on my Mac recently. Getting it to work with coding agents is really easy: install the full Mac app
Neden önemli: 3D sahne ve animasyon üretimini agentic coding ile otomatikleştirmek için doğrudan uygulanabilir bir teknik.
RSS · SIGNAL 4/10

Research acceleration: The view inside OpenAI

Research acceleration: The view inside OpenAI ↗
  • <p><strong><a href="https://openai.com/index/research-acceleration-view-inside-openai/">Research acceleration: The view inside OpenAI</a></strong></p> Apparently today is RSI day at OpenAI, for Recursive Self-Improvement - I think it's their new AGI. Both this piece and the new essay <a href="https:
Neden önemli: OpenAI'nin kendi kendini geliştiren araştırma süreçlerine dair içgörüler, gelecekteki model yeteneklerinin yönünü tahmin etmede faydalı olabilir.
OMNI Labs · otomatik üretildi · kaynak: YouTube transcript + Claude