LIVE · refreshes every 20 min
updated Aug 15, 01:02 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
110010
.art
/ topic
Gemma
Aug 13
1d ago
ggml-cpu/ops vectorizes flash-attention F16 to F32 conversion using F16C intrinsics; PR #26947
r/LocalLLaMA
→ story
1d ago
Muse Glimmer 30B shows potential uses; user asks for clear-cut cases over Gemma 4 31B QAT and Qwen3.6 27B
r/LocalLLaMA
→ story
Aug 12
2d ago
DeepSeek V4 Flash 0731 uncensored jailbreak guide
r/LocalLLaMA
→ story
2d ago
Gemma 4 QAT handles KV cache quantization significantly better, according to KLD benchmarks comparing non-QAT and QAT configurations.
r/LocalLLaMA
→ story
2d ago
Best models for regular machines (16gb ram)
r/LocalLLaMA
→ story
2d ago
Multilingual quantization tax: edge SLMs suffer performance degradation from 4-bit weight quantization across languages, study finds across Gemma 4 and Qwen 3.5 using MMLU ProX Lite and GlobalPIQA
arXiv cs.CL
→ story
Aug 11
3d ago
Gemma and Qwen models may catch hallucinations by examining their own logprobs
r/LocalLLaMA
→ story
3d ago
Mastering edge AI on Raspberry Pi with LiteRT and Gemma
Hacker News (AI)
→ story
3d ago
Gemma 4 E4B and E2B integrated into an e-reader to enable private Q&A and in-app journaling
r/LocalLLaMA
→ story
3d ago
Luth-2: new state-of-the-art French small language models (0.8B and 2.2B) achieve top benchmarks across French tasks
r/LocalLLaMA
→ story
3d ago
Llama-CPP: parallel agents decode well, but one agent’s prefill stalls all others during web searches
r/LocalLLaMA
→ story
Aug 10
4d ago
Ling-3.0-tiny 8B A1.3B MoE released with 1.3B active parameters; Ling team highlights performance between 4B and 8-12B Qwen and Gemma models
r/LocalLLaMA
→ story
4d ago
Chat UIs with native audio input for multimodal models; discussion of direct audio file submission to models without separate STT layer
r/LocalLLaMA
→ story
Aug 9
5d ago
Gemma team to host special event on August 20
r/LocalLLaMA
→ story
6d ago
Qwen 35B A3B tokenizes HTML/JS input differently from Gemma 26B A4B, helping explain why Qwen excels at coding while Gemma performs better on language tasks
r/LocalLLaMA
→ story
Aug 7
7d ago
Why higher micro-batch values may improve llama-cpp model output quality, per user observations
r/LocalLLaMA
→ story
Aug 6
8d ago
Scotoma-2: Gemma4 improved with cleaner writing and fewer tics
r/LocalLLaMA
→ story
8d ago
KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B and Gemma 4 31B; KLD with BeeLlama.cpp v0.4.0 shows KVarN 6-bit beating q8_0 with precision tail dominating
r/LocalLLaMA
→ story
8d ago
10% faster decode with Q4_K MTP draft model using Gemma 4 31b
r/LocalLLaMA
→ story
8d ago
Unsloth's Gemma 4 multimodal features stop working on newer llama.cpp builds; users report broken vision and audio features
r/LocalLLaMA
→ story
8d ago
Artificialanalysis.ai ranks Gemma4 above Qwen3.6 27b in SciCode benchmark results; user questions potential benchmarking issue.
r/LocalLLaMA
→ story
8d ago
The death of SLMs?
r/LocalLLaMA
→ story
8d ago
The Calibration Floor: format repair can masquerade as self-correction at small-to-mid scale
arXiv cs.CL
→ story
Aug 5
9d ago
AttnRes architecture update: core idea remains to replace standard residual stream with attention-based routing between layers; model shows the approach is alive
r/LocalLLaMA
→ story
9d ago
LFM2.5-2.6B runs on OnePlus 13 at ~17 tokens per second on pure CPU; Q4_K_M GGUF used with custom inference engine and device probe suite on Android
r/LocalLLaMA
→ story
Aug 4
10d ago
DeepSeek v4 Flash benchmark compares coding performance against Qwen3.6-27B, 3.5-122B, and Gemma 4 31B using a two-shot file-editing task subset
r/LocalLLaMA
→ story
10d ago
Gemma 4 on 500MB
r/LocalLLaMA
→ story
10d ago
Cost-effective automated judging of natural-language mathematical proofs using open-weight models aligns with human grading decisions on IMO-GradingBench instances
arXiv cs.CL
→ story
Aug 3
11d ago
KAT Coder 2.5 dev: claims faster and more accurate than Qwen 3.6 35b; benchmarks and user impressions cited
r/LocalLLaMA
→ story
11d ago
Knowledge distillation in small instruction-tuned LMs has asymmetric effects on bias, improving context-following for unambiguous tasks but degrading calibration for ambiguous tasks.
arXiv cs.CL
→ story
Aug 2
12d ago
Parlor v2: best-effort fully local GPT-Live clone on an M3 Pro
r/LocalLLaMA
→ story
Aug 1
13d ago
Tomte harness for Gemma 4 released; enables fast local inference on Macs with M processors
r/LocalLLaMA
→ story
13d ago
Are 1B LLMs going away in 2026?
r/LocalLLaMA
→ story
Jul 31
14d ago
Meituan releases LongCat-Flash-Lite-Sparse MoE with ~3B active params and 30B n-gram lookup offloaded to RAM for fast 256k context on 24GB GPU
r/LocalLLaMA
→ story
14d ago
Ported TurboFieldfare to Qwen 3.6 35B; runs in about 1.4 GB RAM
r/LocalLLaMA
→ story
14d ago
Are current LLM benchmarks failing to capture actual usability, as Gemma 4 vs. Gemini/Claude Opus is discussed?
r/LocalLLaMA
→ story
Jul 29
16d ago
Open-source engine runs Gemma 4 26B in 2 GB RAM on any M-series Mac
Hacker News (AI)
→ story
Jul 28
17d ago
Gemma 4 26B/31B Q4 QAT vs Q4/Q5/Q6/Q8
r/LocalLLaMA
→ story
17d ago
Appreciation for Gemma 4 26b A4b
r/LocalLLaMA
→ story
Jul 27
18d ago
Current smallest usable coding model
r/LocalLLaMA
→ story
←
1
2
→