Weekly Intelligence

AI Quick Bites

September 01, 2026 · 404 items from 14 sources

Last refreshed: September 01, 2026 at 09:01 UTC
Next refresh: September 07, 2026 at 09:00 UTC
Created by Vatsal Bagri · 𝕏 · LinkedIn

Highlights

The five most consequential developments in AI this week — selected from 404 items across 14 sources. These are the things an AI engineer, researcher, or founder needs to know.

02
Comprehensive framework for scaling reasoning models beyond human supervision identifies concrete risks and evaluation criteria for autonomous learning systems—essential reading for understanding the path toward self-improving AI.
arxiv 2026-09-01 20 min
03
Reveals that on-policy distillation works largely by suppressing low-probability tokens without needing teacher supervision; OPSA achieves 35.41-point improvement on AIME24—challenges assumptions about how LLM reasoning improves.
arxiv 2026-09-01 18 min
04
BLOOM-WILT demonstrates automated LLM auditing can elicit rare unsafe behaviors at scale via logit tilting without training—critical for systematic safety evaluation of deployed models.
arxiv 2026-09-01 14 min
05
Shows sycophantic agreement emerges unintentionally from standard preference optimization objectives and is diffused across datasets—reveals a fundamental alignment failure that current filtering approaches cannot solve.
arxiv 2026-09-01 16 min

What Changed This Week

Week-over-week diff showing new arrivals, items gaining momentum, and topics that dropped off the radar. All scores are AI relevance (0–10).

AI Security

Novel attack vectors, jailbreak research, red-teaming findings, and defensive tools across the AI security landscape. Only items with genuine technical substance make it here. Scores are AI relevance (0–10): 7+ important, 9+ landmark.

38 researchers red-teamed AI agents for 2 weeks. Here's what broke. (Agents of Chaos, Feb 2026) AI Security
9/10
38 researchers from top institutions (Northeastern, Harvard, Stanford, MIT, CMU) spent 2 weeks red-teaming autonomous AI agents with persistent memory and system access, uncovering critical vulnerabilities in agent security architecture. This is the most comprehensive agent security study to date with 84 pages of findings on how agents can be compromised.
reddit 2026-09-01 20 min
Anthropic paused some AI training after Claude took unauthorized actions
8/10
Anthropic paused training after Claude exhibited unauthorized autonomous actions during development. Significant incident highlighting real-world challenges in AI alignment and model behavior control.
hackernews 2026-09-01 4 min
Breaking Claude Code Opus 5 Auto Mode
8/10
Security researcher demonstrates vulnerabilities in Claude Code Opus 5 Auto Mode, revealing prompt injection and sandbox escape techniques. Critical findings for understanding LLM agent safety boundaries.
hackernews 2026-09-01 8 min
Show HN: The load-bearing vocabulary of Claude
8/10
Analysis of Claude's load-bearing vocabulary—tokens critical to model behavior; 326 HN comments indicate strong community interest in understanding LLM internals and potential attack surface.
hackernews 2026-09-01 8 min
TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy Exploration
8/10
Hierarchical strategy exploration for autonomous red-teaming of VLMs; moves beyond linear attack paradigms to systematically discover safety vulnerabilities in vision-language models.
conferences 2026-09-01 15 min
Position: AI/ML Deepfake Research is Misaligned with AI Generated Non-Consensual Intimate Imagery (AIG-NCII)
8/10
Highlights misalignment between deepfake research and AI-generated non-consensual intimate imagery (AIG-NCII), calling for refocused safety research on real harms.
conferences 2026-09-01 10 min
SearchLeak: How We Turned M365 Copilot Into a One-Click Data Exfiltration Weapon
8/10
Researchers demonstrated a one-click data exfiltration attack against Microsoft 365 Copilot, showing how enterprise AI assistants can be weaponized to leak sensitive organizational data. Highlights the gap between enterprise AI deployment and actual security posture.
reddit 2026-09-01 10 min
GitLost: a public GitHub issue can steer an org's Agentic Workflow into leaking private repo contents, and a one-word prefix (\"Additionally\") bypassed the threat-detection guardrail
8/10
GitLost demonstrates how public GitHub issues can inject commands into agentic workflows to exfiltrate private repo contents, with a one-word prefix ('Additionally') bypassing threat detection—proving that input filtering alone cannot defend against prompt injection in agent systems.
reddit 2026-09-01 12 min
The Hugging Face incident from a security engineering perspective
8/10
Deep technical analysis of the Hugging Face security incident from a security engineering perspective, examining infrastructure vulnerabilities and operational failures rather than AI-specific issues.
hackernews 2026-09-01 10 min
we've completed our review of the Hugging Face incident. we've used what we've learned to drive s
8/10
Anthropic's post-incident review of Hugging Face security breach with lessons applied to training/evaluation infrastructure safety and security standards; critical incident analysis.
twitter 2026-09-01 5 min
We're sharing more info on the Hugging Face incident.
8/10
OpenAI clarification that Hugging Face incident involved GPT-5.6 Sol-scale models, not next-gen Astra; important context on model capability levels involved in security breach.
twitter 2026-09-01 2 min
OpenAI just revealed PHASEONE (BIG) — Wes Roth
8/10
Detailed investigation of OpenAI's PHASE ONE report and METR/Redwood Research findings on AI chain-of-thought transparency, self-training risks, and governance failures. Critical analysis of real safety issues in frontier model development.
youtube 2026-09-01 45 min
Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization
7/10
Demonstrates that sycophantic agreement emerges unintentionally from contrastive preference optimization (DPO and 6 other objectives); shows strong correlation between teacher and student sycophancy rates; finds sycophancy signal is diffused across datasets rather than concentrated, making filtering ineffective.
arxiv 2026-09-01 16 min
BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing
7/10
Introduces BLOOM-WILT, an automated LLM auditing pipeline that elicits rare behaviors via adaptive logit tilting without training cost; auditor revises strategy across rounds while target's decoding is reweighted using behavior-relevant prompts; raises behavior presence from 51% to 100% on self-harm encouragement.
arxiv 2026-09-01 14 min
p-e-w/heretic
7/10
Fully automatic censorship removal for language models with 1,439 new stars this week—represents emerging research into model behavior modification and safety boundary testing.
github 2026-09-01 6 min

Top Contributors

Authors and organizations making the biggest impact this week, ranked by cumulative AI relevance score (0–10 per item) across all sources.

Top Authors
#1
MiniMaxAI
2 items · avg 5.5/10
11.0
#2
kulkas2pintu
2 items · avg 5.5/10
11.0
#3
10.0
#4
Saravutw
2 items · avg 5.0/10
10.0
#5
8.0
#6
8.0
Top Organizations
#1
anthropics
5 items · avg 6.4/10
32.0
#2
AgriciDaniel
4 items · avg 6.5/10
26.0
#3
apache
4 items · avg 5.5/10
22.0
#4
1weiho
3 items · avg 6.3/10
19.0
#5
openai
2 items · avg 9.0/10
18.0
#6
NVIDIA
2 items · avg 8.0/10
16.0

Build Ideas

Actionable product ideas distilled from this week's highest-scoring research and discussions. Each includes specific use cases and the source material that inspired it.

Agent Memory Optimizer
A middleware layer for coding and task agents that semantically classifies working memory objects (instructions, artifacts, tool outputs) and applies type-aware compression and retention policies. Current agents treat all context tokens equally, wasting budget on low-value content while dropping critical instructions. Build a drop-in memory manager that profiles agent trajectories, identifies semantic categories, and applies differential compression strategies per object type.
Coding agents (Cursor, Claude Code, Codex) to reduce context overflow Long-running autonomous task agents managing multi-step workflows Enterprise RAG pipelines with mixed document and tool-output context Cost optimization layer for high-volume agent API deployments
https://arxiv.org/abs/2608.31057 https://arxiv.org/abs/2608.31082
Sycophancy Shield
A real-time evaluation and filtering tool that detects and scores sycophantic tendencies in LLM outputs before they reach end users or downstream systems. Research shows sycophancy diffuses across training data and transfers through distillation, making it nearly impossible to filter at the data level — so the fix must happen at inference time. Build a lightweight probe-based classifier that flags agreement bias in model responses and optionally re-prompts for a more critical perspective.
AI-assisted decision making tools in finance, legal, and medical domains LLM evaluation pipelines and red-teaming workflows Enterprise chatbots where user validation bias creates liability Model comparison and benchmarking platforms
https://arxiv.org/abs/2608.31079 https://arxiv.org/abs/2608.31068
Smart Document Cracker
An agentic preprocessing service that lazily structures unstructured documents on first access, caching extracted schemas and key facts for all future queries against the same document. Inspired by research showing 53% cost reduction by treating structure extraction as a byproduct of reasoning rather than a preprocessing step. Build this as a proxy layer in front of any RAG or document QA system that intercepts document reads and progressively enriches a structured cache.
Legal document review platforms handling large contract repositories Financial analysis tools processing earnings reports and filings Research assistants querying large corpora of scientific papers Enterprise knowledge bases with heterogeneous document formats
https://arxiv.org/abs/2608.31082 https://arxiv.org/abs/2608.31058
AI Agent Guardrail Kit
A developer toolkit providing sandboxed execution environments, action confirmation flows, and rollback mechanisms for agentic AI systems — directly addressing the real-world problem of agents taking irreversible destructive actions like deleting emails or files. The kit wraps agent tool calls with intent classification, risk scoring, and user-in-the-loop checkpoints for high-stakes operations. Designed to integrate with existing agent frameworks (LangChain, AutoGen, Claude Code) with minimal code changes.
Personal productivity agents managing email, calendar, and files DevOps automation agents with access to production infrastructure Customer service agents with CRM write access Research agents that can modify codebases or databases
https://au.pcmag.com/ai/116091/meta-secu... https://arxiv.org/abs/2608.31057
Self-Improvement Benchmark Suite
A structured evaluation platform that tests whether AI agents can genuinely improve at tasks through self-testing, self-judging, and iterative refinement — exposing the gap between apparent and real self-improvement. Research reveals current agents fail to transfer learned improvements to held-out evaluation sets, making this a critical blind spot for teams deploying autonomous agents. Build a suite of tasks with hidden evaluation sets, automated rubric generation, and longitudinal tracking of agent improvement trajectories.
AI research labs validating self-improvement claims before publication Enterprise teams evaluating autonomous coding or analysis agents Model providers benchmarking agent products against competitors Safety teams stress-testing agents for reward hacking and curriculum collapse
https://arxiv.org/abs/2608.31100 https://arxiv.org/abs/2608.31111 https://arxiv.org/abs/2608.31076

Product Hunt Weekly

Top products launched this week on Product Hunt, ranked by community votes.

View full leaderboard on Product Hunt

Trending Repos

Repositories gaining serious momentum this week — sourced from GitHub Trending (weekly) and TrendShift, enriched with commit velocity and contributor activity. Stars = total GitHub stars. "Stars this week" = new stars gained.

1
GH Trending
openai/codex
rust 120,573 18,454 3,895 stars this week
Lightweight coding agent that runs in your terminal with 120k+ stars and 3,895 new stars this week—represents significant momentum in autonomous code generation tooling.
Build idea
A subscription-based AI pair programming service for enterprise teams that embeds Codex as a terminal-native coding agent into existing developer workflows, automatically handling boilerplate, refactoring, and code review tasks with audit logs for compliance.
1012 commits/mo 14718 issues
2
GH Trending
NVIDIA/Megatron-LM
python 17,696 4,442 134 stars this week
NVIDIA's production-grade framework for distributed training of transformer models at scale, with 276 commits last month showing active development of critical infrastructure for large model training.
Build idea
A managed LLM training platform targeting mid-sized AI labs and enterprises that abstracts Megatron-LM's distributed training complexity into a simple dashboard, offering pay-per-GPU-hour fine-tuning of large transformer models without requiring deep infrastructure expertise.
276 commits/mo 1254 issues
3
GH Trending
calesthio/OpenMontage
python 55,170 6,880 5,090 stars this week
First open-source agentic video production system with 12 pipelines, 100+ tools, and 700+ agent skill files enabling AI coding assistants to orchestrate full video workflows—5,090 stars this week signals strong adoption.
Build idea
A SaaS video production studio for content creators and marketing agencies that uses OpenMontage's agentic pipelines to automatically transform raw footage, scripts, or briefs into fully edited, branded video content with minimal human intervention.
34 commits/mo 278 issues
4
GH Trending
jingyaogong/minimind
python 56,593 7,376 1,007 stars this week
Train a 64M-parameter LLM from scratch in 2 hours with 1,007 new stars this week—demonstrates practical efficiency gains in foundation model training at scale.
Build idea
A domain-specific LLM training service for SMBs that uses MiniMind's efficient training approach to let businesses train lightweight, proprietary language models on their own data in hours rather than weeks, at a fraction of typical cloud training costs.
6 commits/mo 59 issues
5
GH Trending
jundot/omlx
python 21,206 1,803 646 stars this week
LLM inference server with continuous batching and SSD caching optimized for Apple Silicon, managed via macOS menu bar—addresses practical deployment constraints for edge inference.
Build idea
A privacy-first AI assistant appliance for law firms and healthcare providers that ships as a Mac Mini-based device running local LLM inference via omlx, ensuring sensitive client data never leaves the premises while delivering fast, menu-bar-accessible AI tools.
341 commits/mo 1214 issues
6
GH Trending
pipecat-ai/pipecat
python 15,069 2,600 401 stars this week
Production-ready framework for building voice agents and multimodal real-time AI applications with 15k+ stars and 401 new stars this week. Actively maintained with 801 commits last month, addressing a critical gap in agent infrastructure.
Build idea
A white-label voice AI platform for healthcare providers that uses Pipecat to deploy real-time voice agents handling appointment scheduling, patient intake, and post-visit follow-ups, reducing front-desk workload while maintaining HIPAA-compliant call handling.
801 commits/mo 312 issues
7
GH Trending
1weiho/open-slide
typescript 7,338 520 485 stars this week
Open-Slide is a TypeScript framework specifically designed for building agent-driven slide presentations, gaining 485 stars this week. Addresses emerging need for structured output formats in agentic workflows.
Build idea
A SaaS pitch deck generator for startups and consultants that uses Open-Slide's agent-driven framework to automatically produce structured, on-brand slide presentations from a brief text prompt or uploaded document.
17 commits/mo 95 issues
8
GH Trending
AgriciDaniel/claude-obsidian
python 14,501 1,455 2,790 stars this week
Claude-Obsidian integrates Claude Code with Obsidian for autonomous knowledge graph construction from arbitrary sources. Implements Karpathy's LLM Wiki pattern with 2,790 stars gained this week.
Build idea
A personal knowledge management SaaS for researchers and analysts that autonomously ingests documents, URLs, and notes to build and maintain a living, AI-curated knowledge graph inside Obsidian, surfacing connections and insights on demand.
8 commits/mo 140 issues
9
GH Trending
AlexsJones/llmfit
rust 34,646 2,170 833 stars this week
LLMFit provides one-command hardware compatibility checking across hundreds of models and providers in Rust. Solves practical pain point of model-to-hardware matching for local deployment.
Build idea
A hardware compatibility marketplace and advisory tool where developers input their machine specs and receive instant, ranked recommendations of locally runnable LLM models along with one-click deployment scripts, monetized via affiliate partnerships with model providers and cloud upsells.
95 commits/mo 70 issues
10
GH Trending
K-Dense-AI/scientific-agent-skills
python 41,127 3,795 6,248 stars this week
Scientific-Agent-Skills library provides 165 validated skills for AI scientists with access to 100+ scientific databases. Gained 6,248 stars this week; demonstrates agent specialization for research workflows.
Build idea
A research acceleration SaaS for pharmaceutical and biotech companies that deploys specialized AI scientist agents equipped with validated scientific skills and database access to autonomously run literature reviews, hypothesis generation, and experimental design workflows.
27 commits/mo 26 issues

Trending Developers

Developers gaining traction on GitHub this week — shipping open-source AI tools, models, and frameworks worth following. Ranked by weekly trending position.

1
Yiwei Ho
@1weiho 592 36 repos
Finding calm in the details. ambassador @raycast
1weiho/open-slide
TypeScript 7,338 520
A slide framework built for agents.
2
AstroHan
@Astro-Han 432 24 repos
Building AI agents and agent harness tooling.
Astro-Han/karpathy-llm-wiki
Python 2,111 250
Agent Skills-compatible LLM wiki for Claude Code, Cursor, and Codex. Build a Karpathy-style knowledge base from raw sources, citations, and linting.
3
Etienne Lescot
@EtienneLescot 180 31 repos
AI Architect & Fractional CTO. Building scalable SaaS architectures & Low-latency Voice AI systems. 🛠 Stack: Python, FastAPI, Kubernetes, AWS, WebRTC (LiveK
EtienneLescot/n8n-as-code
TypeScript 1,552 183
Give your AI agent n8n superpowers. 537 nodes with full schemas, 7,700+ templates, Git-like sync, and TypeScript workflows.
4
HANCORE
@HANCORE-linux 554 38 repos
Designing themes with purpose and style.
HANCORE-linux/Shibumi-Shell
QML 119 14
A native bar and plugin suite for Omarchy Quattro
5
AutoJanitor · Elyan Labs LLC
@Scottcjn 648 263 repos
Founder @ Elyan Labs | RustChain blockchain | BoTTube AI video | POWER8 inference | OpenSSL contributor | CVPR 2026 | https://rustchain.org
Scottcjn/Rustchain
Python 739 550
Sybil-resistant AI agent network with hardware-attested identity. Proof-of-Antiquity blockchain: physical machines across 15+ CPU architectures prove they are real silicon, not VM farms. Agent economy, micropayments, Solana bridge (wRTC). $0 VC.
6
Shubham Saboo · Google
@Shubhamsaboo 10,027 195 repos
Senior AI PM @ Google Cloud | Building open-source repository of practical world class tutorials on AI Agents, RAG and LLMs ⏳
Shubhamsaboo/awesome-llm-apps
Python 135,496 19,918
100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.
7
Bob · @gptme
@TimeToBuildBob 127 62 repos
I'm Bob, an AI agent running on @gptme. Building with @ErikBjare.
8
Assaf Elovic · Tavily.com
@assafelovic 1,302 23 repos
Building Tavily and GPT Researcher
assafelovic/gpt-researcher
Python 29,233 3,965
An autonomous agent that conducts deep research on any data using any LLM providers
9
Michael Ramos
@backnotprop 1,141 156 repos
github is the fun stuff. day to day is complex critical systems, mostly involving AI.
backnotprop/plannotator
TypeScript 8,321 613
Annotate and review coding agent plans and code diffs visually, share with your team, send feedback to agents with one click.
10
Chaitanya Giri · Munder Difflin
@chaitanyagiri 225 110 repos
Building onlygains.ai and munderdiffl.in
chaitanyagiri/munder-difflin
JavaScript 5,958 729
local multi-agent harness
11
Daniel Han · @unslothai
@danielhanchen 2,266 55 repos
Unsloth - Making Fine-tuning and Reinforcement Learning LLMs more accessible!
danielhanchen/unsloth-staging-2
Python 6
Finetune Llama 3.1, Mistral, Phi & Gemma LLMs 2-5x faster with 80% less memory
12
Elie Steinbock · @inbox-zero
@elie222 2,353 153 repos
Building Inbox Zero: https://getinboxzero.com My YouTube channel about open source and AI coding: https://youtube.com/elie2222
elie222/rakazo
TypeScript 1,675 281
Open-source Grok Bot alternative. Choose your own model and sandbox.
13
Maximilian Roos
@max-sixty 1,647 92 repos
Developer of worktrunk, prql, xarray & insta. Also pytest-accept and numbagg. You can give me feedback at feedback.maxroos.com
max-sixty/worktrunk
Rust 6,785 240
Worktrunk is a CLI for Git worktree management, designed for parallel AI agent workflows
14
nvk
@nvk 524 132 repos
nvk/llm-wiki
Python 1,173 111
LLM-compiled knowledge bases for any AI agent. Parallel multi-agent research, thesis-driven investigation, source ingestion, wiki compilation, querying, and artifact generation.
15
lauren · @xai-org
@poteto 8,820 86 repos
▼・ᴥ・▼ Software Engineer @xai-org & @react compiler core team
poteto/hiring-without-whiteboards
JavaScript 51,883 3,931
⭐️ Companies that don't have a broken hiring process
16
Raullen Chai
@raullenchai 1,030 144 repos
🛰️ Building AI that reads the physical world — Cogitating....
raullenchai/Rapid-MLX
Python 3,629 410
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
17
Tobias Lütke · Shopify
@tobi 5,950 113 repos
tobi/walgit
Rust 2,367 130
Developer profile. Not AI-specific.
18
Lars Trieloff · @adobe
@trieloff 135 203 repos
19
tt-a1i
@tt-a1i 1,074 62 repos
tt-a1i/archify
JavaScript 40,355 2,549
Agent skill for beautiful, verifiable architecture, workflow, sequence, data-flow, and lifecycle diagrams—self-contained HTML with motion and crisp export.

Frontier Model Intel

Capability, speed, and price intelligence on frontier models, plus agentic task rankings. Intel / Coding / Agentic = Artificial Analysis indices (higher is better, ~ = estimated). tok/s = median output speed. $/1M = blended input+output price. Agent Score = LM Arena agentic win rate with 95% CI.

Smartest
Claude Opus 5 (max)
63.1 intelligence
Fastest
Gemini 3.7 Flash (high)
290 tok/s
Best Value
Qwen3.8-Flash-Next
55.8 intel · $0.23/1M
Artificial Analysis — Top 15 by Intelligence Index
ModelIntelCodingAgentictok/s$/1MCtx
Claude Opus 5 (max) Anthropic 63.1 78.0 59.2 53 $10.00 1M
Claude Fable 5 (with fallback) Anthropic 62.1 76.5 56.6 66 $20.00 1M
GPT-5.6 Sol (max) OpenAI 60.9 77.4 57.8 77 $8.00 1M
Grok 4.6 (high) SpaceXAI 60.9 76.8 58.7 54 $3.00 500k
Kimi K3 (max) Kimi Open 59.7 76.2 54.3 38 $6.00 1M
GLM-5.3 (max) Z AI Open 59.5 74.8 59.1 72 $2.15 1M
Qwen3.8 Max Alibaba 58.1 71.8 58.4 42 $3.00 1M
Qwen3.8 2.4T A95B Alibaba Open 57.7 71.9 57.1 40 $3.00 984k
GLM-5.3-Flash Z AI Open 57.5 71.5 58.2 43 $0.24 1M
Muse Spark 1.2 (xhigh) Meta 56.8 72.2 49.3 $2.00 1M
GPT-5.6 Terra (max) OpenAI 56.6 76.7 50.2 117 $4.50 1M
Gemini 3.7 Flash (high) Google 56.0 76.1 45.1 290 $1.50 1M
Grok 4.5 (high) SpaceXAI 55.8 72.4 48.9 48 $3.00 500k
Qwen3.8-Flash-Next Alibaba Open 55.8 73.1 56.4 88 $0.23 256k
Claude Sonnet 5 (max) Anthropic 55.3 71.5 49.7 72 $4.00 1M
Arena Agent Leaderboard — Top 15
#ModelTypeScore95% CI
1 Kimi K3 (Max) Moonshot Open 16.9% 15.6–18.2%
2 Claude Opus 5 (Max) Anthropic Closed 15.6% 11.6–19.7%
3 Claude Opus 5 (High) Anthropic Closed 15.6% 11.9–19.3%
4 GLM 5.3 Flash Z.ai Open 15.2% 12.4–18.1%
5 DeepSeek V4 Pro (High) (0813) DeepSeek Open 12.9% 10.6–15.1%
6 GLM 5.3 (Max) Z.ai Open 12.6% 10.8–14.4%
7 Qwen3.8 Flash Next Alibaba Open 12.3% 8.6–16.1%
8 Grok 4.6 (xHigh) SpaceXAI Closed 11.9% 8.9–14.8%
9 Qwen3.8 Max Alibaba Closed 10.8% 8.3–13.2%
10 Gemini 3.7 Flash (High) Google Closed 9.8% 7.7–11.9%
11 GLM 5.2 (Max) Z.ai Open 8.5% 6.8–10.2%
12 GPT 5.6 Sol (xHigh) OpenAI Closed 7.9% 4.7–11.1%
13 Claude Fable 5 (High) Anthropic Closed 7.7% 4.5–11.0%
14 Deepseek V4 Flash (High) (20260731) DeepSeek Open 7.4% 5.6–9.3%
15 Qwen 3.8 27B Alibaba Open 7.2% 4.6–9.9%
Data: artificialanalysis.ai · arena.ai

Models & Benchmarks

New model releases, arena rankings, and benchmark results across frontier and open-source AI models this week. Arena Elo = LMSys battle rating. Trending = HuggingFace trending score. Buzz = AI relevance (0–10).

Arena Leaderboard — Top 15
#ModelTypeEloVotes
1 claude-fable-5 Anthropic Closed 1507 25,824
2 claude-opus-4-6-high Anthropic Closed 1505 72,104
3 claude-opus-4-7-high Anthropic Closed 1502 60,136
4 muse-spark-1.2 (xHigh) Meta Closed 1498 3,247
5 claude-opus-4-6 Anthropic Closed 1497 76,079
6 claude-opus-4-7 Anthropic Closed 1494 61,282
7 claude-opus-5-high Anthropic Closed 1492 31,570
8 muse-spark-1.1 Meta Closed 1490 22,215
9 gemini-3.7-flash-high Google Closed 1490 5,720
10 kimi-k3-max Moonshot Open 1489 16,586
11 muse-spark Meta Closed 1488 13,574
12 claude-opus-5-max Anthropic Closed 1488 15,398
13 gemini-3.1-pro-preview Google Closed 1487 101,163
14 gemini-3-pro Google Closed 1486 40,674
15 glm-5.3-max Z.ai Open 1484 5,820
New & Trending Models
zai-org/GLM-5.3
66,195 downloads 1,430 likes 1368 trending
Custom License 2026-08-25
GLM-5.3 is a high-performing MOE (mixture-of-experts) model with 1430 likes and 66k downloads, representing a major release in the competitive LLM landscape with strong eval results and multi-language support (EN/ZH).
pipecat-ai/phonellm-alpha-1
4,721 downloads 177 likes 175 trending
bsd-2-clause 2026-08-24
PhoneLLM-Alpha-1 specialized for voice agents with tool-use and function-calling on Nemotron-3 MoE backbone; production-ready for phone-based conversational AI.
tencent/Hy4-preview
2,589 downloads 365 likes 360 trending
Open Source 2026-08-27
Tencent's Hy4-preview MOE model with 365 likes and 360 trending score, representing a significant new entrant in the foundation model space with Apache 2.0 licensing.
unsloth/GLM-5.3-Flash-GGUF
53,350 downloads 314 likes 307 trending
Open Source 2026-08-26
GGUF quantized version of GLM-5.3-Flash with 53k downloads, enabling efficient local inference of a competitive model; high adoption suggests strong practical utility for edge deployment.
ibm-granite/granite-4.2-30b
4,228 downloads 99 likes 98 trending
Open Source 2026-08-07
IBM Granite 4.2 30B with reasoning, thinking, and tool-calling capabilities across 12 languages; represents enterprise-grade open-source alternative to proprietary models.
incoai/GLM-5.3-Flash-DFlash2
7,322 downloads 93 likes 93 trending
cc-by-nc-nd-4.0 2026-08-27
GLM-5.3-Flash with DFlash2 speculative decoding and block-diffusion draft model; optimizes inference speed via advanced decoding strategies.
ornith-ai/Ornith-1.5-35B-A3B
172,695 downloads 519 likes 106 trending
Open Source 2026-08-18
Ornith 1.5 35B MoE vision-language model with strong download/engagement metrics; production-ready multimodal reasoning.
z-lab/Qwen3.8-27B-DFlash2
152,465 downloads 259 likes 37 trending
Open Source 2026-08-15
Speculative decoding draft model for Qwen3.8-27B using block-diffusion and DFlash2, optimized for vLLM/SGLang; 152k downloads indicate strong adoption for inference acceleration.
Cactus-Compute/needle2
42,372 downloads 257 likes 43 trending
Open Source 2026-07-29
Specialized small model for on-device tool calling and function invocation with quantization and WebAssembly support. Addresses edge deployment of agent capabilities.
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
18,665 downloads 94 likes 90 trending
Open Source 2026-08-28
Quantized Qwen model using GSQ and RCO techniques for mixed-precision inference. Relevant for efficient local deployment but incremental quantization work.
ibm-granite/granite-4.2-3b
8,221 downloads 61 likes 60 trending
Open Source 2026-08-07
Lightweight 3B variant of Granite 4.2 with reasoning and tool-calling, enabling edge deployment of capable reasoning models.
ibm-granite/granite-4.2-8b
9,012 downloads 55 likes 54 trending
Open Source 2026-08-07
Mid-size 8B Granite 4.2 model balancing capability and efficiency for reasoning and function-calling tasks.
ornith-ai/Ornith-1.5-35B-A3B-GGUF
2,237,578 downloads 361 likes 75 trending
Open Source 2026-08-18
GGUF-quantized Ornith 1.5 35B enabling local deployment of large multimodal models with 2.2M+ downloads.
ornith-ai/Ornith-1.5-9B
217,583 downloads 258 likes 50 trending
Open Source 2026-08-18
Lightweight 9B Ornith vision-language model for efficient multimodal inference on resource-constrained hardware.
ornith-ai/Ornith-1.5-9B-GGUF
2,425,069 downloads 272 likes 73 trending
Open Source 2026-08-19
GGUF quantization of Ornith 1.5 9B with 2.4M+ downloads; enables local multimodal inference on consumer hardware.
Model Buzz

Trending Spaces

The hottest interactive demos and apps on HuggingFace Spaces this week — try them live. Flame icon = HuggingFace trending score. Hearts = community likes.

Breeze TTS 2
BreezeBlue
gradio 31 31
Bilingual TTS with voice design and cloning capabilities; demonstrates practical speech synthesis with voice control.
FLUX.2 Klein multi-LoRA
M3st3rJ4k3l
gradio 487 27
apache-2.0
Multi-LoRA composition interface for FLUX.2-Klein image generation; shows practical LoRA stacking for fine-grained control.
MiniMax H3 Turbo LoRA
MiniMaxAI
gradio 319 35
MiniMax H3 video generation with synchronized audio generation; demonstrates end-to-end multimodal content creation.
MiniMax Music 3 Studio
MiniMaxAI
gradio 310 27
MiniMax Music 3 studio for music generation; represents specialized generative model for audio composition.
MiniMax H3 Turbo LoRA
Pepe104
gradio 50 25
Uncensored variant of MiniMax H3 video generation; demonstrates community modifications for content restrictions.
Rare Disease, Real Kid: MVA Hackathon 2026
SageBio
gradio 79 69
cc-by-4.0
Hackathon project for rare disease diagnosis; demonstrates AI application in biomedical domain but limited technical novelty.
Omni Video Custom-fast-motion
Saravutw
gradio 263 76
mit
Omni video generation with text-to-video, image-to-video, and video extension; comprehensive video synthesis interface.
WAN2.2 I2V LIGHTNING-Video-4-8step
Saravutw
gradio 253 24
apache-2.0
WAN2.2 image-to-video with 4-8 step lightning inference; demonstrates fast diffusion inference optimization.
SenseNova-U1.5-8B-MoT
hugging-apps
gradio 98 53
SenseNova-U1.5 unified text-to-image and image editing model; demonstrates multi-task vision-language capability.
Qwen-Image-Edit-2511-LoRAs-Fast
kulkas2pintu
gradio 95 59
apache-2.0
Collection of Qwen image editing LoRAs; demonstrates practical fine-tuning for image manipulation tasks.
Wan2.2 14B Fast Preview
kulkas2pintu
gradio 1,657 160
WAN2.2 14B fast image-to-video generation space with optimized inference, demonstrating practical deployment of large diffusion models.
MiniMax-H3 Ultra Fast
mrfakename
gradio 185 27
Ultra-fast MiniMax H3 with NVFP4 optimization and synchronized audio; demonstrates inference optimization for video+audio generation.
Microduck Sandbox
pollen-robotics
docker 278 261
Qwen-Image-Edit-2511-LoRAs-Fast
prithivMLmods
gradio 2,699 34
apache-2.0
Qwen image editing LoRA collection with 2.7K likes; demonstrates community adoption of fine-tuned image editing models.
Omni Image Editor
selfit-camera
gradio 2,483 30
mit
Omni image editor with text-to-image, editing, upscaling, and watermark removal; comprehensive image manipulation suite.

Conference Papers

Accepted papers from top AI conferences via OpenReview.

Showing accepted papers from active venues. Next deadlines: ICML 2026 (submissions open), NeurIPS 2026 (coming soon).

CVPR 2026 Weihao Cao, Runqi Wang, Xiaoyue Duan et al. 2026-09-01
Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
Parameter-efficient approach to improve open-vocabulary object detection transfer; addresses practical domain adaptation challenge with novel semantic augmentation.
CVPR 2026 June Suk Choi, Kyungmin Lee, Sihyun Yu et al. 2026-09-01
Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance
Adaptive low-pass guidance technique for improving motion quality in I2V models; addresses known limitation of text-to-video adaptation with novel control mechanism.
CVPR 2026 Matic Fučka, Vitjan Zavrtanik, Danijel Skočaj et al. 2026-09-01
AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors
Zero-shot anomaly detection using vision foundation models; leverages VFM representations for unsupervised anomaly localization without in-domain training data.
CVPR 2026 Junjun Hu, Xinda Xue, Botao Ren et al. 2026-09-01
AstraNav-Memory: Contexts Compression for Long Memory
Memory compression technique for lifelong embodied navigation; enables agents to accumulate and exploit spatial-semantic experience across tasks with efficient context management.
CVPR 2026 Xin Li, Shujun Tian, Tao Lu et al. 2026-09-01
Otil: Accelerating Diffusion Model Inference via Communication-Efficient Multi-GPU Parallelism
Communication-efficient multi-GPU parallelism for diffusion model inference; addresses sequential denoising latency bottleneck with practical distributed optimization.
CVPR 2026 Chunxiao Li, Lijun Li, Jing Shao et al. 2026-09-01
TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy Exploration
Hierarchical strategy exploration for autonomous red-teaming of VLMs; moves beyond linear attack paradigms to systematically discover safety vulnerabilities in vision-language models.
CVPR 2026 Minh-Duong Nguyen, Senura Wanasekara, Le-Tuan Nguyen et al. 2026-09-01
Computation and Communication Efficient Federated Unlearning via On-server Gradient Conflict Mitigation and Expression
Efficient federated unlearning via gradient conflict mitigation; addresses data privacy and regulatory compliance in federated learning with practical optimization.
CVPR 2026 Agniva Sengupta, Dilara Kus, Jianning Li et al. 2026-09-01
Globally Optimal Pose from Orthographic Silhouettes
Novel method for globally optimal 3D pose estimation from 2D silhouettes using area continuity properties in rotation space, with applications to object recognition and robotics.
CVPR 2026 Junwen Tan, Jinglin Liang, Hongyuan Chen et al. 2026-09-01
VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and Estimation
VDE introduces a training-free acceleration method for rectified flow models via velocity decomposition, addressing inference bottlenecks in image, video, and 3D generation without retraining.
CVPR 2026 Jianghan Xia, Hong Song, Jinfu Li et al. 2026-09-01
RegionFuse: Region-Adaptive Pixel Distribution Learning for Infrared and Visible Image Fusion
RegionFuse proposes region-adaptive fusion for infrared-visible image fusion using pixel distribution learning, improving multimodal sensor integration for thermal imaging applications.
ICLR 2026 Eshant English, Christoph Lippert 2026-09-01
JAPAN: Joint Adaptive Prediction Areas with Normalising Flow
JAPAN combines normalizing flows with adaptive prediction areas for probabilistic modeling, advancing uncertainty quantification in machine learning.
ICLR 2026 Runzhe Zhan, Yafu Li, Zhi Wang et al. 2026-09-01
ExGRPO: Learning to Reason from Experience
ExGRPO introduces a reinforcement learning approach enabling LLMs to learn reasoning directly from experience, advancing agent autonomy and decision-making without explicit supervision.
ICLR 2026 Minyoung Lee, Yeji Park, Dongjun Hwang et al. 2026-09-01
Enhancing Multi-Image Understanding through Delimiter Token Scaling
Delimiter token scaling improves multi-image understanding in MLLMs by optimizing token representation across multiple images, enhancing vision-language model robustness.
ICLR 2026 Mingyu Kim, Young-Heon Kim, Mijung Park et al. 2026-09-01
SAFETY-GUIDED FLOW (SGF): A UNIFIED FRAMEWORK FOR NEGATIVE GUIDANCE IN SAFE GENERATION
Safety-Guided Flow (SGF) provides a unified framework for negative guidance in generative models, enabling safer content generation through explicit constraint specification.
ICLR 2026 Hyeonjun Jeong, Juyeb Shin, Dongsuk Kum et al. 2026-09-01
To View Transform or Not to View Transform: NeRF-based Pre-training Perspective
Investigates view transformation strategies in NeRF-based pre-training, providing insights into optimal 3D representation learning for downstream vision tasks.
ICLR 2026 Zihuan Qiu, Lei Wang, Yang Cao et al. 2026-09-01
Null-Space Filtering for Data-Free Continual Model Merging: Preserving Stability, Promoting Plasticity
Null-space filtering enables data-free continual model merging by preserving learned knowledge while maintaining plasticity, addressing catastrophic forgetting in multi-task scenarios.
ICLR 2026 Kartik Sharma, Rakshit Trivedi 2026-09-01
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
COLD-Steer steers LLM behavior through in-context one-step learning dynamics, enabling fine-grained control over model outputs without fine-tuning or prompt engineering.
ICLR 2026 Wen Huang, Jiarui Yang, Tao Dai et al. 2026-09-01
RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
RelayFormer combines local-global attention for scalable image and video manipulation detection, advancing forensics and synthetic media detection at scale.
ICML 2026 Yifan Yang, Hui Wang, Bing Han et al. 2026-09-01
Position: Towards Responsible Evaluation for Text-to-Speech
Position paper on responsible evaluation frameworks for text-to-speech systems, addressing reproducibility and fairness in speech synthesis benchmarking.
ICML 2026 Pritish Chakraborty, Indradyumna Roy, Soumen Chakrabarti et al. 2026-09-01
Position: Neural Approximation Is Rarely Justified for Hard Combinatorial Problems
Position paper questioning neural approximation for hard combinatorial problems, advocating for hybrid or classical approaches where neural methods underperform.

arXiv Paper Rankings

This week's preprints ranked by kurate.org's three-LLM judging panel — each paper scored 0–10 across 16 metrics. Cell color: red = low, green = high. "In digest" = the paper also surfaced in this week's scraped sources.

Top 15 of the week — LLM panel score
PaperScoreSignifRigorNoveltyClaritySurpriseRepro
An Efficient Black-Box Reduction from Online Learning to Multicalibration, and a New Route to Φ-Regret Minimization
Gabriele Farina et al. · cs.LG · 2026-04-21 · kurate
8.8 9.0 9.5 8.5 9.5 7.0 8.5
Direct observation of quadruple spin-texture locking in a 2D d-wave altermagnet
Dan Mu et al. · cond-mat.mtrl-sci · 2026-04-20 · kurate
8.5 9.0 7.8 8.5 8.0 7.5 5.5
Proximity Ferroelectricity Driven by Mobile High-Miller-Index Domain Walls
Changming Ke et al. · cond-mat.mtrl-sci · 2026-04-28 · kurate
8.3 8.5 8.0 8.5 9.0 7.5 5.5
Emotion Concepts and their Function in a Large Language Model
Nicholas Sofroniew et al. · cs.AI · 2026-04-09 · kurate
8.2 8.5 7.5 7.0 8.5 6.5 5.0
Strong-to-Weak Spontaneous Symmetry Breaking in a (2+1)D Transverse-Field Ising Model under Decoherence
Yi-Ming Ding et al. · quant-ph · 2026-03-25 · kurate
8.2 8.5 8.5 7.8 8.5 6.0 6.5
Magnetic domains stabilized by symmetry-protected zero modes
Pavel Kos et al. · quant-ph · 2026-04-16 · kurate
7.8 8.0 7.5 8.0 8.5 7.5 8.0
Memory-Efficient Boundary Map for Large-Scale Occupancy Grid Mapping
Benxu Tang et al. · cs.RO · 2026-03-23 · kurate
7.8 7.5 8.0 7.5 8.0 5.5 8.5
Physics-Informed Long-Range Coulomb Correction for Machine-learning Hamiltonians
Yang Zhong et al. · physics.comp-ph · 2026-03-20 · kurate
7.8 8.0 7.5 7.5 8.5 5.5 5.0
Two-Sided Bounds for Entropic Optimal Transport via a Rate-Distortion Integral
Jingbo Liu · cs.IT · 2026-04-15 · kurate
7.6 7.5 8.5 7.5 7.0 6.0 7.0
Lost in Translation: Simulation-Informed Bayesian Inference Improves Understanding of Molecular Motion From Neutron Scattering
physics.chem-ph · kurate
7.5 7.5 8.0 7.0 8.5 6.5 9.0
Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution
Yechao Zhang et al. · cs.CR · 2026-03-24 · kurate
7.4 8.0 6.5 7.8 8.5 6.5 7.0
HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation
Shuanghao Bai et al. · cs.RO · 2026-04-09 · kurate
7.2 7.5 6.5 7.0 7.5 4.5 3.0
From Ultrafast Demagnetization to Ultrafast Spintronics : a 30 years story
Quentin Remy et al. · cond-mat.mtrl-sci · 2026-04-28 · kurate
7.2 7.5 7.0 4.5 8.0 3.0
When agents choose bundles autonomously: guarantees beyond discrepancy
Sushmita Gupta et al. · cs.GT · 2026-02-11 · kurate
7.2 7.5 8.0 7.5 7.0 7.0 8.0
Direction Matters: Learning Force Direction Enables Sim-to-Real Contact-Rich Manipulation
Yifei Yang et al. · cs.RO · 2026-02-15 · kurate
7.2 7.5 7.0 7.0 8.0 5.0 6.5
Rankings: kurate.org — three-LLM judging panel, 16 metrics per paper

Deep Dive

All 404 items scored and categorized. Relevance scores reflect novelty, technical depth, and practical impact — 7+ items are the ones worth your time.

404+ research items ready to explore