LIVE · refreshes every 20 min
updated Sep 3, 19:22 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
110010
.art
/ topic
Qwen
Aug 14
20d ago
Qwen 30b MoE runs at 30 tps on 6 GB VRAM with Hermes 90k context; user reports 20–35 tps during actual session generation
r/LocalLLaMA
→ story
20d ago
1BIT Qwen 3.8 2.4T a95b: user reports 508 GB model usage, 50pp and 9.6 tgen with latency drop at 50k tokens; notes Unslloth Studio on Mac Studio Ultra performing.
r/LocalLLaMA
→ story
Aug 13
20d ago
DS4 cloud vs. Qwen3.6 36B vs. Muse Glimmer 30B speed benchmark on Llama.cpp (RTX 5080)
r/LocalLLaMA
→ story
20d ago
Waiting for Qwen 3.8 27B evokes Star Wars Episode I disappointment; user shares hype ahead of release
r/LocalLLaMA
→ story
20d ago
Qwen 3.8 release adds prompt-steered reasoning effort; fix discusses issues with Jinja chat template, including inability to disable thinking and blank think tags
r/LocalLLaMA
→ story
21d ago
Qwen 3.8Max 2.4t Open Weight lacks vision; user expresses disappointment over size and lack of vision capability
r/LocalLLaMA
→ story
Aug 12
21d ago
Multi-model workflows show benefit from planner-actor style with larger models; user compares Qwen 27b token usage against simpler prompting
r/LocalLLaMA
→ story
21d ago
Qwen 3.8 27B: will it use DFlash or MTP head?
r/LocalLLaMA
→ story
22d ago
Qwen 3.8 2.4T released; 27B not released today
r/LocalLLaMA
→ story
22d ago
Qwen set to launch in just over seven hours; Google Translate not ready, according to the excerpt.
r/LocalLLaMA
→ story
22d ago
Cracks in the foundation: four minor architectural decisions reduce long context extensibility across Olmo, Llama, and Qwen dense models
arXiv cs.CL
→ story
22d ago
Multilingual quantization tax: edge SLMs suffer performance degradation from 4-bit weight quantization across languages, study finds across Gemma 4 and Qwen 3.5 using MMLU ProX Lite and GlobalPIQA
arXiv cs.CL
→ story
22d ago
Tinkerer shares personal upgrades for local agents, including switching to llama.cpp and exploring harnesses like hermes agent and pi.dev
r/LocalLLaMA
→ story
22d ago
27B Q8 and 35B Q6 tested on a 32 GB GPU; neither model proved reliable enough to serve as its own final checker
r/LocalLLaMA
→ story
Aug 11
22d ago
CMP170HX performance tested on mining cards; can run multiple small models on a single 64GB card without exhaustively testing 8B/12B models
r/LocalLLaMA
→ story
22d ago
Gemma and Qwen models may catch hallucinations by examining their own logprobs
r/LocalLLaMA
→ story
23d ago
Unsloth Desktop app released for Mac, Windows, and Linux to run and train models locally, with support for MLX, diffusion models, audio models, and GGUF, plus local Claude/Code integration and sandboxed execution.
r/LocalLLaMA
→ story
23d ago
Small open-weight AI models threaten AI development, says author
r/LocalLLaMA
→ story
23d ago
12GB VRAM gang, what’s our plan? focuses on qwen finetuned MoEs and dense models for smaller setups; asks whether upgrading to 24GB VRAM is the only option
r/LocalLLaMA
→ story
23d ago
Qwen 3.8-27b coming this week
r/LocalLLaMA
→ story
23d ago
Interpreting reasoning mechanisms of large language models via sparse autoencoders: separating Thinking from NoThinking in CoT-enabled models
arXiv cs.CL
→ story
Aug 10
23d ago
Ling 3.0 Flash on Strix Halo; vLLM ROCm/HiP, int4 tensors, Qwen-122b on rocmFP4 compared for speed (not a fair comparison)
r/LocalLLaMA
→ story
24d ago
Best open-source alternative to Claude Code for local models with 1:1 interface like Claude Code
r/LocalLLaMA
→ story
24d ago
Ling-3.0-tiny 8B A1.3B MoE released with 1.3B active parameters; Ling team highlights performance between 4B and 8-12B Qwen and Gemma models
r/LocalLLaMA
→ story
24d ago
Best current ERP base model that are smart and uncensored?
r/LocalLLaMA
→ story
24d ago
Motif-Technologies advances to next round in South Korea's AI Foundation Model project, edging out Qwen 3.7 Max and LG EXAONE; Upstage's Solar Pro 4 exp leads Motif's score.
r/LocalLLaMA
→ story
24d ago
1M-context with a 17 GB model on 24 GB VRAM achieved on RTX 3090; user manu69x reports stable 1M context with Qwen 3.5 35B A3B and extraction of 7 needles from text
r/LocalLLaMA
→ story
24d ago
Comparing how Cline, Kilo, and Qwen Code handle long-task context and why context loops recur
r/LocalLLaMA
→ story
24d ago
Running Qwen 3.5 35B A3B-Q8_0 gguf on a Radeon 7600 at 18 token/s with 64 GB DDR4 RAM and Ryzen 5600 using llama.cpp on Ubuntu
r/LocalLLaMA
→ story
Aug 9
24d ago
endless-frontier/BigBang-v1 - qwen 3.5 finetunes benchmark results; comparisons between DeepSeek Flash and Pro show per-benchmark variance
r/LocalLLaMA
→ story
25d ago
AMD patch reduces MTP buffer overhead, increasing Qwen 27B context from 64K to 149K
r/LocalLLaMA
→ story
25d ago
Apple says Mac users in China can connect to Alibaba's Qwen AI service
Hacker News (AI)
→ story
25d ago
Show HN: Qwen3.8-Max uses Qwen Studio and MCP to code locally for free
Hacker News (AI)
→ story
25d ago
Qwen 35B A3B tokenizes HTML/JS input differently from Gemma 26B A4B, helping explain why Qwen excels at coding while Gemma performs better on language tasks
r/LocalLLaMA
→ story
Aug 8
26d ago
I built a local realtime voice stack for Ollama: Parakeet STT → Qwen 2.5 7B → Qwen3-TTS
r/LocalLLaMA
→ story
26d ago
Qwen 3.8 gains attention; user shares experience with 3.6 27B and predicts at-home LLM serving could bypass metered AI services
r/LocalLLaMA
→ story
26d ago
Qwen 35B-A3B MoE is ~4× faster than Qwen 27B dense on local coding tests with a smaller quality gap than expected
r/LocalLLaMA
→ story
Aug 7
26d ago
Qwen 3.6 27B flags/settings in llama.cpp
r/LocalLLaMA
→ story
27d ago
Is GLM 5.2 and Kimi 2.7 still relevant amid newer models like Kimi K3, Qwen 3.8 Max, and V4 pro Deepseek?
r/LocalLLaMA
→ story
27d ago
Artificial Analysis updates its intelligence index to adjust rankings; user alleges bias and paid-off influence
r/LocalLLaMA
→ story
←
1
2
3
4
5
6
→