LIVE · refreshes every 20 min
updated Sep 3, 19:22 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
110010
.art
/ topic
Qwen
Aug 6
27d ago
Dual RTX 3090s and ~50 GB RAM used to run DSV4; user prepares to test a reap model and hopes for performance better than Qwen 27B
r/LocalLLaMA
→ story
28d ago
Qwen 3.8 Max ranked best overall model ahead of Opus 5 by Artificial Analysis agentic index
r/LocalLLaMA
→ story
28d ago
KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B and Gemma 4 31B; KLD with BeeLlama.cpp v0.4.0 shows KVarN 6-bit beating q8_0 with precision tail dominating
r/LocalLLaMA
→ story
28d ago
Two 5070 Ti GPUs run Qwen 27B full config with cu129-nightly KV cache improvements; 2 concurrent threads maintain speed, decode at 94-87 tps, 0-120k context, prefill 4.6k-2.4k, 170k GPU KV plus 246k with 8GB RAM
r/LocalLLaMA
→ story
28d ago
Non-Coding Harness Terminal UI prompts for general-purpose TUI use; user seeks recommendations before selecting a terminal UI solution
r/LocalLLaMA
→ story
28d ago
The death of SLMs?
r/LocalLLaMA
→ story
28d ago
Frontier narrows pricing gap as performance catches up to price concerns; users report downgraded free version after new models release
r/LocalLLaMA
→ story
28d ago
Qwen3.8-2.4T-A95B (aka Qwen3.8-Max) open release next Wednesday
r/LocalLLaMA
→ story
28d ago
GLM 5.2 outperforms Qwen3.6 27b in coding tasks and API usage; user prefers GLM 5.2 for its reliability and ease of use
r/LocalLLaMA
→ story
Aug 5
29d ago
LFM2.5-2.6B runs on OnePlus 13 at ~17 tokens per second on pure CPU; Q4_K_M GGUF used with custom inference engine and device probe suite on Android
r/LocalLLaMA
→ story
29d ago
TensorSharp adds MoE CPU-offload feature to main, enabling up to 35B-A3B MoE to fit with long-context KV cache on 12-16 GB GPUs; router, shared experts stay on accelerator.
r/LocalLLaMA
→ story
29d ago
Qwen Developers outline plans for 27B and 122B models; AMA responses suggest more models may arrive and crit pit scores are unavailable in cards
r/LocalLLaMA
→ story
29d ago
SLMs and QAT: labs debate training precision versus efficient quantization and the push to smaller models like Nanbeige, Liquid, and Qwen; LFM2.5 2.6B released and planned for Q8 testing
r/LocalLLaMA
→ story
29d ago
Building a fully local pdf read-aloud and pdf-to-audiobook desktop app with Kokoro 82M, Qwen, and llama.cpp
r/LocalLLaMA
→ story
Aug 4
29d ago
GPT-OSS turns one year old; user shares favorable comparisons to Qwen 3.5 122B and Nemotron 3 Super in local models
r/LocalLLaMA
→ story
29d ago
Qwen 3.8 Max improves over Qwen 3.7 Max on the Debate Benchmark (1462 → 1588); average cost per debate increases by 45%
r/LocalLLaMA
→ story
29d ago
Local LLM 35B MoE shows real-world coding benchmarks; Qwen 3.6, Ornith, and KAT compared in practical dev workflows
r/LocalLLaMA
→ story
Aug 4, 2026
Mach-1 Additive achieves 95% of Qwen 3.6 35B performance while being 10x smaller
r/LocalLLaMA
→ story
Aug 4, 2026
Company approves 128GB Mac for research proposal on running local LLMs; Qwen 3.6/3.8 considered best in 20–60GB space, is there improvement with higher memory or stronger model?
r/LocalLLaMA
→ story
Aug 4, 2026
Decreasing the power limit of the 5090 to 480W yields negligible inference slowdown, per test with Qwen 3.6-27b.
r/LocalLLaMA
→ story
Aug 3
Aug 3, 2026
Hugging Face artificial analysis performs well for general work; coding is the only area inferior to Qwen.
r/LocalLLaMA
→ story
Aug 3, 2026
Qwen says next week 3.8 will be open weights
r/LocalLLaMA
→ story
Aug 3, 2026
Was the release of deepseek v4 flash planned to take spotlight against 5.6 luna?
r/LocalLLaMA
→ story
Aug 3, 2026
Deep dive on OPD and RL for LLMs
r/LocalLLaMA
→ story
Aug 3, 2026
KAT Coder 2.5 dev: claims faster and more accurate than Qwen 3.6 35b; benchmarks and user impressions cited
r/LocalLLaMA
→ story
Aug 3, 2026
Qwen will release new models every month
r/LocalLLaMA
→ story
Aug 2
Aug 2, 2026
Quantizing KV Cache for DeepSeek V4 Flash degrades quality, per DS4F perplexity results
r/LocalLLaMA
→ story
Aug 2, 2026
All Qwen model oneshots: 1109 outputs across 33 models and 35 prompts to compare
r/LocalLLaMA
→ story
Aug 2, 2026
Best model around 3B parameters for multilingual understanding and instruction following?
r/LocalLLaMA
→ story
Aug 2, 2026
Qwen 3.8 has dropped; day 90 status update shows no release yet
r/LocalLLaMA
→ story
Aug 1
Aug 1, 2026
Are 1B LLMs going away in 2026?
r/LocalLLaMA
→ story
Aug 1, 2026
Qwen 3.6 27B Q5 runs on 3x2080ti with 55 tps using llama.cpp; user asks if more performance can be squeezed.
r/LocalLLaMA
→ story
Jul 31
Jul 31, 2026
SenseNova releases U1.5 Lite preview with benchmarks and improvements in 4K native generation, text rendering, and structured prompts; benchmarks show gains across ImgEdit-Bench, Qwen-Image-Bench, and GEdit-Bench-en.
r/LocalLLaMA
→ story
Jul 31, 2026
Meituan releases LongCat-Flash-Lite-Sparse MoE with ~3B active params and 30B n-gram lookup offloaded to RAM for fast 256k context on 24GB GPU
r/LocalLLaMA
→ story
Jul 31, 2026
What’s the real value of high reasoning in AI models?
r/LocalLLaMA
→ story
Jul 31, 2026
Ported TurboFieldfare to Qwen 3.6 35B; runs in about 1.4 GB RAM
r/LocalLLaMA
→ story
Jul 31, 2026
Can we expect Deepseek v4 distills into smaller models?
r/LocalLLaMA
→ story
Jul 31, 2026
Benchmarking the big-model orchestrator plus local-model worker split: latency and data transfer concerns reduce the win
r/LocalLLaMA
→ story
Jul 31, 2026
Optimal realistic local AI for most
r/LocalLLaMA
→ story
Jul 30
Jul 30, 2026
Would extremely high decode tok/s be useful for inference?
r/LocalLLaMA
→ story
←
1
2
3
4
5
6
→