LIVE · refreshes every 20 min
updated Sep 3, 18:23 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
110010
.art
/ topic
GPT-5
Sep 2
1d ago
Do multimodal LLMs see before they read? Diagnosing contextual sycophancy in multimodal reasoning
arXiv cs.CL
→ story
Sep 1
2d ago
GeoJSON Map Viewer
Simon Willison
→ story
2d ago
Thinking costs tokens: adding inference structure hurts performance below a token-budget threshold and helps above it
arXiv cs.AI
→ story
Aug 27
7d ago
Pushing GPT-5.6 Luna from 0% to 56% on ARC-AGI-3 Public.
Hacker News (AI)
→ story
Aug 19
15d ago
Replit expands access to software creation with GPT-5.6 Luna, introducing Free Mode
OpenAI
→ story
15d ago
Do all your agents need models like Claude 5 or GPT-5.6?
Hacker News (AI)
→ story
Aug 17
16d ago
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index, tying with GPT-5.6 Luna Max and trailing GLM-5.2 Max and DeepSeek V4 Pro 0813 Max
Simon Willison
→ story
16d ago
GPT-5.6 Sol reduces pricing by 50%
Hacker News (AI)
→ story
Aug 15
19d ago
CORS Chat provides a web UI to test an OpenAI-Responses-compatible chat endpoint; conversations persist in the browser and can be exported as JSON.
Simon Willison
→ story
Aug 14
20d ago
Show HN: APIMart aggregates discounted AI APIs for GPT-5 and Sora 2
Hacker News (AI)
→ story
Aug 13
21d ago
The builder’s guide to GPT‑5.6; Startups use GPT‑5.6 to build faster, more cost-efficient AI agents with smarter model selection and new Responses API capabilities.
OpenAI
→ story
21d ago
Preview Ultrafast: OpenAI offers GPT-5.6 Sol up to 14× faster with Cerebras-powered service
OpenAI
→ story
Aug 12
21d ago
Alchemy-utils 0.1a0 release introduces a database-agnostic library prototype inspired by sqlite-utils, focusing on core API methods like insert, upsert, insert_all, upsert_all, create, update, and table introspection
Simon Willison
→ story
Aug 11
23d ago
GPT-5.6-Sol Pro achieves 67.30% on Claude’s Reimann paper evaluation
Hacker News (AI)
→ story
Aug 10
24d ago
Model ML uses GPT-5.6 Sol to carry finance work from research and analysis through editable, traceable PowerPoint decks and Excel workbooks.
OpenAI
→ story
24d ago
OpenAI releases GPT-5.6-Cyber through Daybreak Red for authorized vulnerability research and security testing
OpenAI
→ story
Aug 7
27d ago
Universal Pathologies, Conditional Consequences: A triple-robustness analysis of RAG for multi-hop traceability
arXiv cs.CL
→ story
Aug 2
Aug 2, 2026
June newsletter: highlights include accidental cyberattacks by OpenAI and Anthropic models under test, GPT-5.6 Sol/Terra/Luna, Claude Opus 5, Kimi K3 and DeepSeek-V4-Flash-0731, and other model releases
Simon Willison
→ story
Jul 30
Jul 30, 2026
llm 0.32rc2 release fixes dependency issue and updates default model to GPT-5.6 Luna; users can switch back to 4o mini via llm models default gpt-4o-mini
Simon Willison
→ story
Jul 30, 2026
OpenAI makes frontier models cheaper with GPT-5/6 price-performance improvements
r/LocalLLaMA
→ story
Jul 29
Jul 29, 2026
Two API settings tripled scores on the ARC-AGI-3 benchmark by GPT-5.6, boosting performance and efficiency through retained reasoning and enabled compaction.
OpenAI
→ story
Jul 29, 2026
GPT-5.6 vs. Claude Fable 5 for physical AI: which performs best?
Hacker News (AI)
→ story
Jul 29, 2026
Personalization, personas, and forecasting in value alignment
arXiv cs.AI
→ story
Jul 29, 2026
GPT-5.6 improves AI efficiency across models, inference, and agentic workflows to deliver more useful intelligence per dollar.
OpenAI
→ story
Jul 26
Jul 26, 2026
Harness showdown: Claude Code, OpenCode, and Pi produce similar quality on DeepSeek V4 Flash benchmark; Claude Code takes longer despite faster tokens in some configurations
r/LocalLLaMA
→ story
Jul 25
Jul 25, 2026
How much are you actually using your local models these days; which ones do you reach for the most?
r/LocalLLaMA
→ story
Jul 23
Jul 23, 2026
Structured synthetic reasoning data improves arithmetic fine-tuning for small language models, using a 21,250-example GPT-5-mini-generated corpus derived from GSM8K
arXiv cs.AI
→ story
Jul 22
Jul 22, 2026
GPT-5.6 got smarter; then it kept acting.
Hacker News (AI)
→ story
Jul 22, 2026
We probed a pinned GPT-5.5 endpoint; every request carried about 1,447 hidden tokens
Hacker News (AI)
→ story
Jul 22, 2026
Relay-Bench evaluates LLMs on multi-domain reasoning chains; GPT-5.5 (xHigh) scores 43.3% on composite problems across domains
arXiv cs.CL
→ story
Jul 21
Jul 21, 2026
Drawing the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
Hacker News (AI)
→ story
Jul 20
Jul 20, 2026
GPT-5.6 Sol and Kimi K3 compete in Kerbal Space Program speedrunning live event
Hacker News (AI)
→ story
Jul 20, 2026
Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment
arXiv cs.CL
→ story
Jul 18
Jul 18, 2026
Claude Fable 5 to be included in all Max and Team Premium plans; Pro and Team Standard keep access via credits with a $100 one-time credit
Simon Willison
→ story
Jul 17
Jul 17, 2026
GPT-5.6 Sol Ultra constructed Chrome V8 exploit chain from patch commits
Hacker News (AI)
→ story
Jul 17, 2026
GPT-5.6 Sol Max released; analysts assess its value
Hacker News (AI)
→ story
Jul 16
Jul 16, 2026
Gpt-5.6 Sol Pro solves open problem in convex optimization.
Hacker News (AI)
→ story
Jul 16, 2026
Moonshot AI unveils Kimi K3, a 2.8-trillion-parameter model described as their most capable to date, with open weights promised by July 27, 2026.
Simon Willison
→ story
Jul 16, 2026
Datasette code-frequency chart on GitHub shows spikes in activity aligned with Opus 4.8, GPT-5.5, Fable 5 and GPT-5.6 Sol.
Simon Willison
→ story
Jul 16, 2026
DOOMQL uses SQLite as the game engine for a Python terminal Doom-like game; project by Peter Gostev built with GPT-5.6 Sol
Simon Willison
→ story
←
1
2
→