LIVE · refreshes every 20 min
updated Aug 26, 19:42 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
110010
.art
/ topic
Llama
Aug 5
21d ago
Qwen3-TTS voice cloning merged into mainline llama.cpp with full support for speaker references and multi-language generation
r/LocalLLaMA
→ story
Aug 4
21d ago
Liquid AI releases LFM2.5-2.6B with 128K context and tool calling, runs at 30 tokens per second on a phone
r/LocalLLaMA
→ story
22d ago
Llama.cpp adds GPU-based sampling for MTP, claiming about an 8% speedup in tok/s on Qwen3.6-35B with a 5090; observed ~4% speedup on Nvidia P40 in tests.
r/LocalLLaMA
→ story
22d ago
Migrate from LM Studio to llama.cpp; user experiences and learning curve discussed
r/LocalLLaMA
→ story
22d ago
New DSv4 Flash Doom Loop in Q8? Llama.cpp Vulkan
r/LocalLLaMA
→ story
22d ago
Researchers discuss creating a website to share hardware specs with specific llama.cpp flags to identify what works.
r/LocalLLaMA
→ story
Aug 3
22d ago
NousResearch continues development on Hermes and releases Hermes agent 0.20; project launched at 0.2 in mid-March.
r/LocalLLaMA
→ story
23d ago
Qwen3-Next runs at full speed with MTP support in ggml-org/llama.cpp pull request 25589
r/LocalLLaMA
→ story
23d ago
Strategies for capping thinking on ds4 flash 0731
r/LocalLLaMA
→ story
Aug 2
23d ago
Llama.app Mac app and llama serve from llama.cpp released by the llama.cpp team
r/LocalLLaMA
→ story
24d ago
TPS halved after long use with llama.cpp setup on 3070; user reports sudden drop in throughput with same command and models
r/LocalLLaMA
→ story
24d ago
LocalLLaMA remains a strong source of open-weight research, but locating it requires navigating benchmark discussions and hardware-focused posts on the subreddit.
r/LocalLLaMA
→ story
24d ago
Network inference using llama.cpp rpc-server across four systems with multiple GPUs and over 10000 version; experiments benchmark MoE models locally
r/LocalLLaMA
→ story
24d ago
How well do multiple GPUs scale for LLM inference?
r/LocalLLaMA
→ story
Aug 1
25d ago
Fix for Deep Seek v4 Flash 0731 tool calling added to llama.cpp
r/LocalLLaMA
→ story
25d ago
Acquarium Panel Failure on DS4 flash 0731; Q3_K_XL Unsloth | DS4 flash 0731 issue reported with Acquarium panel failure on Q3_K_XL Unsloth
r/LocalLLaMA
→ story
25d ago
Are 1B LLMs going away in 2026?
r/LocalLLaMA
→ story
25d ago
Qwen 3.6 27B Q5 runs on 3x2080ti with 55 tps using llama.cpp; user asks if more performance can be squeezed.
r/LocalLLaMA
→ story
Jul 31
26d ago
Uncensored multi-model releases include LongCat-Flash-Lite with MTPs, Jamba2-Mini, Qwen3.5-9B-Nikusui-v1 with MTPs, and Qwen3.5-27B-Nikusui-v1 with MTPs, available in safetensors and GGUF formats
r/LocalLLaMA
→ story
26d ago
Can we expect Deepseek v4 distills into smaller models?
r/LocalLLaMA
→ story
Jul 30
27d ago
Inkling-Small by thinkingmachines: 276B total parameters, 12B active, 1M context window.
r/LocalLLaMA
→ story
27d ago
How close are we to local llama robotics for consumer price point?
r/LocalLLaMA
→ story
27d ago
Nanbeige4.2-3B underperforms vs. Qwen3.5-9B and Gemma4-12B in tests, despite promised benchmarks; user aims for a fast coding-focused model to replace Qwen3.6-35B
r/LocalLLaMA
→ story
27d ago
Does MTP head load in VRAM by default?
r/LocalLLaMA
→ story
27d ago
Mechanistic interpretability streamlined for everyday users; open-source tool supports local LLMs like GPT-2 and Llama for basic to expert analysis
r/LocalLLaMA
→ story
27d ago
Benchmark results for MindControl for llama.cpp released, including HumanEval+ and LiveCodeBench benchmarks
r/LocalLLaMA
→ story
27d ago
Micro-Llama training on a published dataset fails to run in Unsloth Studio; user reports issues using ShortChildrenStories dataset
r/LocalLLaMA
→ story
Jul 29
28d ago
5060ti Chads, vllm updates and nvfp4
r/LocalLLaMA
→ story
28d ago
model: add NextN/MTP speculative decoding support for GLM_DSA (GLM-5.2)
r/LocalLLaMA
→ story
28d ago
GBNF grammar compiler enables 8B models to reliably call tools; details of the approach and implementation provided in the deep dive
r/LocalLLaMA
→ story
28d ago
TimeCapsule: a 1.2B LLaMA-style model trained only on Victorian texts (1800–1875) as an epistemologically isolated generative archive.
arXiv cs.CL
→ story
28d ago
Show HN: I run a 30B model at 22 tokens/s and 109 tokens/s, not novel; 6GB/16GB RAM usage with llama.cpp
Hacker News (AI)
→ story
Jul 28
29d ago
Llama.cpp updates include chunked SSD matmul acceleration for Mamba-2 prefill and a ggml-metal FWHT kernel for the metal backend
r/LocalLLaMA
→ story
29d ago
Spec: add DSpark speculative decoding in llama.cpp PR 25173 by ggml-org
r/LocalLLaMA
→ story
29d ago
Update chat template for dsv4 in llama.cpp to override gguf’s template with new DeepSeek-V4.jinja
r/LocalLLaMA
→ story
29d ago
HDL emerges as a hard decision layer in transformers, causing abrupt stabilization of answer option rankings during inference
arXiv cs.AI
→ story
29d ago
Semalith v1.4, a 184M DeBERTa-v3-base classifier, achieves state-of-the-art prompt-injection detection with 44x fewer parameters than Llama-Guard-3-8B
arXiv cs.LG
→ story
Jul 27
Jul 27, 2026
Ling-3.0-flash: SGLang commits to day-0 support; vLLM says support coming soon; llama.cpp marks 2.6 request as not_planned
r/LocalLLaMA
→ story
Jul 27, 2026
[Model] Add support for Nanbeige4.2 by zqlcode in llama.cpp pull request #25994
r/LocalLLaMA
→ story
Jul 27, 2026
Current smallest usable coding model
r/LocalLLaMA
→ story
←
1
2
3
4
→