AI Digest
28 September 2026 · 5 kaynak
⚠ Kaynak uyarisi
- Higgsfield Blog (rss) · 5 gundur sessiz
Feed bozulmus olabilir: site yeniden yapilmis, URL 404 donuyor, ya da gonderen adresi degismis olabilir. Kontrol et.
RSS · SIGNAL 6/10
2026 in LLMs (so far)
2026 in LLMs (so far) ↗- <p>On Friday I gave the closing keynote at the <a href="https://www.wearedevelopers.com/world-congress-north-america">WeAreDevelopers World Congress North America</a> in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026.
Neden önemli: Simon Willison'ın 2026 LLM trendlerine dair genel bakışı, agentic coding ve model gelişmelerini takip etmek isteyen bir stüdyo için faydalı olabilir.
NEWSLETTER · SIGNAL 6/10
TypeSafe AI's Jev launches a new 'System One' model category for fast, cheap structured decisions—triggering rapid open-source alternatives like Laya and Stanford/Nvidia's CLM-8B.
🚀 System One models are carving out a new layer in the AI stack ↗- TypeSafe AI released Jev on Sept 15: a non-generative model returning typed decisions (e.g. 'billing' vs text) in 70-500ms at $0.042/M input tokens with no output-token cost.
- Jev claims 40-200x lower latency and 444.6x cost reduction vs frontier LLMs for classification/routing/scoring tasks.
- Open-source alternative Laya (421M params, ModernBERT-based) offers a Jev-compatible API and ~33ms inference on a Tesla T4.
- Stanford/Nvidia's CLM-8B uses contrastive embeddings instead of classification heads, running up to 9x faster than Jev in zero-shot tests but sometimes with lower accuracy (CLM: 81.6%/87.6% vs Jev: 71.1%/83.1% on coding-judge benchmarks).
- Emerging pattern: 'Jev-as-judge' replaces expensive LLM-as-judge—generative models produce candidates, cheap decision models rank/route them, escalating only low-confidence cases to LLMs or humans.
- Trained via RLCD (Reinforcement Learning for Calibrated Decisions) so confidence scores are meaningful—0.52 vs 0.998 can trigger different handling logic.
Neden önemli: AI creative studio'nuz icin bu, uretim hatlarindaki 'fuzzy karar' katmanlarini (icerik moderasyonu, agent tool secimi, aday cikti siralama, kalite kontrol) pahali LLM cagrilariyla degil ucuz System One modelleriyle cozebileceginiz anlamina geliyor. Ozellikle coklu-aday uretim + secim workflow'larinizda (ornegin en iyi gorseli/metni secme), Jev veya Laya gibi modelleri 'judge' olarak kullanarak maliyeti dusurup hizi artirabilirsiniz—ama secim uzayinin (kategori/secenekler) eksiksiz ve net tanimlanmasi gerekiyor, yoksa model yanlis secim yapabilir.
NEWSLETTER · SIGNAL 6/10
Crusoe's new Serverless Fine-Tuning platform publishes full benchmarks and costs, showing small fine-tuned models can beat giant ones at a fraction of the price.
⚙️ Crusoe brings receipts to AI fine-tuning ↗- Crusoe's Serverless Fine-Tuning is now GA, covering 19 open models from 2B to 750B parameters, each tested with a public, reproducible benchmark harness.
- Fine-tuning GLM 5.2 (753B params) cost just $180 and boosted intent classification by 19pp and multi-turn conversation scores by 61pp.
- A fine-tuned Qwen3.5-2B outperformed Qwen3-235B on the Banking77 customer-service benchmark at roughly 1/15th the training cost.
- Pricing is fully transparent and tiered: $0.40 per million tokens under 16B params up to $10 above 300B, no sales calls required.
- Crusoe openly reported regressions (Qwen3.5-9B and Gemma-4-31B got worse on a theorem-proving benchmark after fine-tuning), flagging likely pre-training contamination instead of hiding it.
- Fine-tuned models can deploy straight to dedicated endpoints via Self-Serve Deployments starting at $5.50/hour on Nvidia H100s.
Neden önemli: Bir AI yaratici studyosu icin en kritik nokta: kucuk, dogru fine-tune edilmis bir model, cok daha buyuk ve pahali bir modelden hem daha iyi sonuc verebiliyor hem de maliyeti onlarca kat dusurebiliyor - bu da ozel is akislarina yatirim yapmayi mantikli kiliyor. Ayrica seffaf benchmark ve fiyatlandirma, hangi modelin gercekten ise yaradigini 'satis vaadine' guvenmeden kendi verinizle test edebilmenizi sagliyor.
NEWSLETTER · SIGNAL 4/10
CrowdStrike's product chief says AI agents don't just attack faster, they chain tactics into novel attack patterns that break human-reasoning-based defenses.
⚙️ Cybersecurity’s entry-level job is changing ↗- AJ Shipley (CrowdStrike CPO) argues AI agents 'reason differently' than humans, combining known TTPs into unfamiliar attack chains that 20-30 year old security architectures weren't built to detect
- Entry-level SOC jobs are shifting up-stack: AI now handles tier-1 alert triage and initial evidence-gathering, so 'entry-level' will mean tier-2/tier-3 work instead
- Shipley frames this as a productivity multiplier, not headcount replacement — same analyst count expected to do 10x-100x more work with AI tooling
- CrowdStrike's proposed fix for AI non-determinism: ensemble of models 'voting' on outcomes plus input-controlling harnesses, similar to old multi-engine malware sandboxing
- 3-5 year bull case: 'red models' auto-discover vulnerabilities, 'blue models' auto-remediate them across endpoint/cloud/SaaS/network before exploitation
- Sponsor segment (Coder) pushes concept of a dedicated 'AI Operating Layer' sitting outside the app/data/compute stack for governance and visibility
Neden önemli: Bu haber doğrudan yaratıcı AI stüdyosu işiyle ilgili değil ama iki transfer edilebilir ders var: iş katmanlarının yukarı kayması (rutin işler AI'ya devrolduğunda 'giriş seviyesi' tanımı değişiyor) ve non-deterministik AI çıktılarını ensemble/harness yaklaşımıyla kontrol etme fikri — bu, ajan tabanlı yaratıcı pipeline'lar kurarken de işe yarayabilir. Sponsorlu 'AI Operating Layer' kavramı ise altyapı/governance tarafında izlenmeye değer bir trend.
RSS · SIGNAL 4/10
Kākāpō Party
Kākāpō Party ↗- <p><strong>Tool:</strong> <a href="https://tools.simonwillison.net/kakapo-party">Kākāpō Party</a></p> <p>I presented a closing keynote for the <a href="https://www.wearedevelopers.com/world-congress-north-america">WeAreDevelopers World Congress North America</a> yesterday. As <a href="https://simonw
Neden önemli: Simon Willison'ın yeni aracı, yaratıcı iş akışlarına entegre edilebilecek ilginç bir prototip sunuyor, ancak doğrudan üretim değeri sınırlı.
OMNI Labs · otomatik üretildi · kaynak: YouTube transcript + Claude