AI Digest
14 September 2026 · 3 kaynak
⚠ Kaynak uyarisi
- Higgsfield Blog (rss) · 4 gundur sessiz
- HuggingFace Blog (rss) · 4 gundur sessiz
Feed bozulmus olabilir: site yeniden yapilmis, URL 404 donuyor, ya da gonderen adresi degismis olabilir. Kontrol et.
NEWSLETTER · SIGNAL 7/10
DeepSeek-V4.1-Flash doubles model size while cutting KV cache memory 4x through a redesigned encoder-decoder and sparse attention architecture.
⚡ What DeepSeek-V4.1-Flash teaches us about efficient AI ↗- DeepSeek-V4.1-Flash is a 552B MoE model (8B active params for input, 16B for output) using a Causal Encoder-Decoder split (20 encoder + 20 decoder layers) so prompt tokens skip the full backbone.
- KV cache footprint dropped from 3,514 bytes/token (V4-Flash) to 890 bytes/token, cutting HBM needs to 1/4 and SSD cache storage to 1/8 of the predecessor.
- Sliding-Window Attention with 'Bounded Replay' lets the model cheaply reconstruct local context after a session pause, avoiding costly full recomputation.
- New Compressed Sparse Attention 2 (CSA2) lets attention layers share KV state and retrieval results via Full/Reindex/Reuse layer modes, cutting redundant memory across layers.
- A Hierarchical Sparse Indexer bounds search cost for 1M-token contexts by narrowing candidates from 2,048 blocks down to 512 positions, regardless of total context length.
- On Artificial Analysis Intelligence Index, V4.1-Flash scores 40 (vs Gemini 3.8 Flash High's 41) at roughly 1/4 the cost per task ($0.30/M input, $1.20/M output tokens).
Neden önemli: Uzun context'li ajanlar veya coklu tool-call zincirleri kuran bir stüdyo icin bu, parametre sayisinin artik maliyet tahmininde yeterli olmadigi anlamina geliyor — asil soru KV cache/token ve oturum devam ettirme maliyeti. DeepSeek'in mimari yaklasimi (encoder-decoder ayrimi, sparse attention, hiyerarsik indeksleme) inference maliyetlerini dusurmek icin model kucultmekten daha etkili bir yol oldugunu gosteriyor; bu teknikler yayilirsa uzun-context agent workloadlari icin fiyatlandirma beklentilerini degistirebilir.
RSS · SIGNAL 6/10
Generating running routes with GPT-6 Astra and ChatGPT Work
Generating running routes with GPT-6 Astra and ChatGPT Work ↗- <p>Here's a neat thing I had <a href="https://simonwillison.net/2026/Aug/30/understanding-chatgpt-work/">ChatGPT Work</a> with GPT-6 Astra (Max) do this morning:</p> <blockquote> <p><code>I live at <my address>. Figure out 5K and 10K running routes from me that loop from my house. Use OSM data
Neden önemli: GPT-6 Astra'nın agentic görev yürütme yeteneğini gösteriyor, workflow otomasyonu için ilham verebilir.
RSS · SIGNAL 4/10
shot-scraper 1.12
shot-scraper 1.12 ↗- <p><strong>Release:</strong> <a href="https://github.com/simonw/shot-scraper/releases/tag/1.12">shot-scraper 1.12</a></p> <p>I've added WebP support to my <a href="https://shot-scraper.datasette.io/">shot-scraper</a> screenshot automation tool. You can now take a WebP screenshot of a web page like t
Neden önemli: Ekran görüntüsü alma aracına WebP desteği eklenmesi, görsel üretim iş akışlarında dosya boyutunu optimize etmek için işine yarayabilir.
OMNI Labs · otomatik üretildi · kaynak: YouTube transcript + Claude