LIVE · refreshes every 20 min
updated Aug 26, 19:42 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
110010
.art
/ topic
Llama
Aug 12
14d ago
Tinkerer shares personal upgrades for local agents, including switching to llama.cpp and exploring harnesses like hermes agent and pi.dev
r/LocalLLaMA
→ story
Aug 11
15d ago
Ling-3.0 support added to llama.cpp in a 40-line PR for Tiny model; PR 26608 still not merged to mainline
r/LocalLLaMA
→ story
15d ago
Are third-party gguf/mmproj safe on Llama in production environments?
Hacker News (AI)
→ story
15d ago
Built a low-power llama.cpp server with an Intel N100 and RTX 5060Ti
r/LocalLLaMA
→ story
15d ago
Add CI targets for ROCm 7.14 in llama.cpp; PR 25775 adds Linux and Windows targets for ROCm 7.14 support
r/LocalLLaMA
→ story
15d ago
Llama-CPP: parallel agents decode well, but one agent’s prefill stalls all others during web searches
r/LocalLLaMA
→ story
Aug 10
15d ago
Muse Spark 1.2 Open Source before Llama 4 Behemoth discussed; Muse AI community notes open sourcing Muse Spark 1.2 while Llama 4 Behemoth awaits release
r/LocalLLaMA
→ story
16d ago
I compared GGUF quants of Qwen3.6 27B to NVFP4, AWQ, AutoRound, and FP8.
r/LocalLLaMA
→ story
16d ago
DiffusionGemma technical report published on arXiv
r/LocalLLaMA
→ story
16d ago
Ante 0.2 enables a self-contained coding agent that manages llama.cpp offline by pointing at a GGUF; it handles the inference engine, supported on Metal, CUDA, Vulkan, or CPU, with pinned builds and model discovery.
r/LocalLLaMA
→ story
16d ago
Chat UIs with native audio input for multimodal models; discussion of direct audio file submission to models without separate STT layer
r/LocalLLaMA
→ story
16d ago
Running Qwen 3.5 35B A3B-Q8_0 gguf on a Radeon 7600 at 18 token/s with 64 GB DDR4 RAM and Ryzen 5600 using llama.cpp on Ubuntu
r/LocalLLaMA
→ story
16d ago
WinterMix 59 GiB Qwen3.5-122B-A10B build with 20k+ context achieves improved perplexity via new reasoning trace annealing; MLX on Apple Silicon faster than llama.cpp on same hardware, Apache 2.0 weights on HF
r/LocalLLaMA
→ story
Aug 9
16d ago
KLQ: Training-free measured rotation quantization outperforms other training-free rotation-based methods on W4A4KV4-bits; Llama 3.2 1B KLQ-quantized beats SpinQuant and nears ReSpinQuant without GPTQ/LDLQ rounding
r/LocalLLaMA
→ story
17d ago
Dual Radeon AI PRO R9700 server slower than RTX 5090 for LLM inference; Ollama bottleneck, vLLM/llama.cpp and other recommendations?
r/LocalLLaMA
→ story
17d ago
Underestimated budget solution: Ryzen 7 260/Ryzen 9 8945HX with 780m iGPU and 64 GB RAM cited as affordable option for GPU-accelerated tasks
r/LocalLLaMA
→ story
17d ago
AMD patch reduces MTP buffer overhead, increasing Qwen 27B context from 64K to 149K
r/LocalLLaMA
→ story
Aug 8
17d ago
DSpark draft model runs extremely slowly (1-2 t/s) with DeepSeek-V4-Flash on llama-server vs MTP
r/LocalLLaMA
→ story
17d ago
Enabling PCIe peer-to-peer for consumer Nvidia cards yields more performance than expected
r/LocalLLaMA
→ story
18d ago
Local 4x 6000 Pro (multi-year progression) showcased with a multi-year image series of a local AI cluster including 4x RTX 6000 Pro Max Q and 4x 3090s to run locally and keep data out of the cloud
r/LocalLLaMA
→ story
18d ago
llamacpp runs slower than Ollama on a Huihui-Qwen3.6-35B-A3B-abliterated-ggml-model-Q4_K gguf setup, 55 tps vs. 61 tps for Ollama
r/LocalLLaMA
→ story
18d ago
Tesla V100 users sought to share config and performance for Qwen3.6 27B on Q4_K_M + Q8_0 MTP 128K context
r/LocalLLaMA
→ story
18d ago
Qwen3.5-9B MTP GGUF shows 35.8% acceptance on CUDA and 91–92% on Vulkan; two Radeon/RADV systems are 77% and 128% faster on Vulkan.
r/LocalLLaMA
→ story
18d ago
ggml-org/llama.cpp adds Longcat-Flash support for testing; PR 19182 seeks tests on larger models with GGUF excerpt from HuggingFace
r/LocalLLaMA
→ story
18d ago
Qwen 35B-A3B MoE is ~4× faster than Qwen 27B dense on local coding tests with a smaller quality gap than expected
r/LocalLLaMA
→ story
18d ago
GFX1030/Multiple V620 users: set -ub 384 and -b to a multiple of that to fix llama.cpp tensor split crashes
r/LocalLLaMA
→ story
18d ago
PR speeds 300GB model loads from 4m54s to 1m38s; GGML_RPC_LOAD_THREADS optimization on 4060Ti platforms reduces load time to ~1.5 minutes
r/LocalLLaMA
→ story
18d ago
BeeLLama issues with kvarn6 cache types and llama-server options for Qwen3.6-35B-A3B-IQ4 model
r/LocalLLaMA
→ story
Aug 7
18d ago
Qwen 3.6 27B flags/settings in llama.cpp
r/LocalLLaMA
→ story
19d ago
Self-taught professional becomes Director of AI and Systems development; shares journey from indie game dev to leveraging LLMs like Vicuna and LLaMA for AI work
r/LocalLLaMA
→ story
19d ago
Why higher micro-batch values may improve llama-cpp model output quality, per user observations
r/LocalLLaMA
→ story
19d ago
llama.cpp PR enables x86 VNNI for Q2_0; 8B decode throughput rises from 2.39 to 8.20 tok/s on Bonsai GGUFs
r/LocalLLaMA
→ story
19d ago
Echo Dot 2 runs 28M LLMs locally with llama.cpp and offline speech recognition
r/LocalLLaMA
→ story
Aug 6
19d ago
AMD acquires Taalas; startup demos Llama 3.1 8B at 17k tok/s
Hacker News (AI)
→ story
20d ago
KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B and Gemma 4 31B; KLD with BeeLlama.cpp v0.4.0 shows KVarN 6-bit beating q8_0 with precision tail dominating
r/LocalLLaMA
→ story
20d ago
Unsloth's Gemma 4 multimodal features stop working on newer llama.cpp builds; users report broken vision and audio features
r/LocalLLaMA
→ story
20d ago
Output-budget regimes change the measured multilingual reasoning gap in MGSM for Qwen3-8B and Llama-3.1-8B-Instruct across prompting budgets
arXiv cs.CL
→ story
Aug 5
21d ago
Using an Nvidia RTX 5090 with 64 GB RAM and an older AMD GPU with 12 GB VRAM to run a local Llama.cpp model; a trade business asks if both GPUs can be utilized for two different AI models.
r/LocalLLaMA
→ story
21d ago
TensorSharp adds MoE CPU-offload feature to main, enabling up to 35B-A3B MoE to fit with long-context KV cache on 12-16 GB GPUs; router, shared experts stay on accelerator.
r/LocalLLaMA
→ story
21d ago
Building a fully local pdf read-aloud and pdf-to-audiobook desktop app with Kokoro 82M, Qwen, and llama.cpp
r/LocalLLaMA
→ story
←
1
2
3
4
→