LIVE · refreshes every 20 min
updated Sep 3, 20:02 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
110010
.art
/ topic
Gemma
Aug 19
15d ago
Prompt extend model fine-tuned from gemma-4-12B-it for Qwen Image Edit 2511 generates enhanced editing prompts using PERL with ROLL; Kimi K2.6 serves as reward worker to evaluate results
r/LocalLLaMA
→ story
Aug 18
16d ago
Agent on CPU, which to pick?
r/LocalLLaMA
→ story
Aug 17
17d ago
Opus 4.8, Muse Glimmer, Gemma 4 draw criticism for reasoning behavior; user complaints discussed in thread
r/LocalLLaMA
→ story
17d ago
Ling 3.0 Tiny 8b runs fastest on low-end PC with 4GB VRAM, 36 tokens/sec; claims strong performance vs Qwen 3.5 9b and Gemma 12
r/LocalLLaMA
→ story
17d ago
Mimir: a 1.7B model claims to outperform Qwen 3.5 0.8B and Gemma 4 E2B on benchmarks (english and danish; builds on sapient’s hrm-text)
r/LocalLLaMA
→ story
17d ago
Local models demonstrate SVG generation across Qwen3.8-27B, Muse Glimmer 30B, Gemma 4 26B A4B, and Gemini 3.7 Flash as control.
r/LocalLLaMA
→ story
Aug 15
18d ago
Google could release a 120B dense multimodal Gemma model to challenge OAI and Anthropic, says submitter.
r/LocalLLaMA
→ story
19d ago
Gemma 4 E4B IQ2_XXS tensor level allocation recovers reasoning performance from 28.9 to 69.5 at the same 3.3 GB budget
r/LocalLLaMA
→ story
Aug 13
21d ago
ggml-cpu/ops vectorizes flash-attention F16 to F32 conversion using F16C intrinsics; PR #26947
r/LocalLLaMA
→ story
21d ago
Muse Glimmer 30B: user asks for clear-cut standout use cases vs. Gemma 4 31B QAT and Qwen3.6 27B
r/LocalLLaMA
→ story
Aug 12
22d ago
DeepSeek V4 Flash 0731 uncensored jailbreak guide
r/LocalLLaMA
→ story
22d ago
Gemma 4 QAT handles KV cache quantization significantly better, according to KLD benchmarks comparing non-QAT and QAT configurations.
r/LocalLLaMA
→ story
22d ago
Best models for regular machines (16gb ram)
r/LocalLLaMA
→ story
22d ago
Multilingual quantization tax: edge SLMs suffer performance degradation from 4-bit weight quantization across languages, study finds across Gemma 4 and Qwen 3.5 using MMLU ProX Lite and GlobalPIQA
arXiv cs.CL
→ story
Aug 11
22d ago
Gemma and Qwen models may catch hallucinations by examining their own logprobs
r/LocalLLaMA
→ story
23d ago
Mastering edge AI on Raspberry Pi with LiteRT and Gemma
Hacker News (AI)
→ story
23d ago
Gemma 4 E4B and E2B integrated into an e-reader to enable private Q&A and in-app journaling
r/LocalLLaMA
→ story
23d ago
Luth-2: new state-of-the-art French small language models (0.8B and 2.2B) achieve top benchmarks across French tasks
r/LocalLLaMA
→ story
23d ago
Llama-CPP: parallel agents decode well, but one agent’s prefill stalls all others during web searches
r/LocalLLaMA
→ story
Aug 10
24d ago
Ling-3.0-tiny 8B A1.3B MoE released with 1.3B active parameters; Ling team highlights performance between 4B and 8-12B Qwen and Gemma models
r/LocalLLaMA
→ story
24d ago
Chat UIs with native audio input for multimodal models; discussion of direct audio file submission to models without separate STT layer
r/LocalLLaMA
→ story
Aug 9
24d ago
Gemma team to host special event on August 20
r/LocalLLaMA
→ story
25d ago
Qwen 35B A3B tokenizes HTML/JS input differently from Gemma 26B A4B, helping explain why Qwen excels at coding while Gemma performs better on language tasks
r/LocalLLaMA
→ story
Aug 7
27d ago
Why higher micro-batch values may improve llama-cpp model output quality, per user observations
r/LocalLLaMA
→ story
Aug 6
27d ago
Scotoma-2: Gemma4 improved with cleaner writing and fewer tics
r/LocalLLaMA
→ story
28d ago
KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B and Gemma 4 31B; KLD with BeeLlama.cpp v0.4.0 shows KVarN 6-bit beating q8_0 with precision tail dominating
r/LocalLLaMA
→ story
28d ago
10% faster decode with Q4_K MTP draft model using Gemma 4 31b
r/LocalLLaMA
→ story
28d ago
Unsloth's Gemma 4 multimodal features stop working on newer llama.cpp builds; users report broken vision and audio features
r/LocalLLaMA
→ story
28d ago
Artificialanalysis.ai ranks Gemma4 above Qwen3.6 27b in SciCode benchmark results; user questions potential benchmarking issue.
r/LocalLLaMA
→ story
28d ago
The death of SLMs?
r/LocalLLaMA
→ story
28d ago
The Calibration Floor: format repair can masquerade as self-correction at small-to-mid scale
arXiv cs.CL
→ story
Aug 5
29d ago
AttnRes architecture update: core idea remains to replace standard residual stream with attention-based routing between layers; model shows the approach is alive
r/LocalLLaMA
→ story
29d ago
LFM2.5-2.6B runs on OnePlus 13 at ~17 tokens per second on pure CPU; Q4_K_M GGUF used with custom inference engine and device probe suite on Android
r/LocalLLaMA
→ story
Aug 4
Aug 4, 2026
DeepSeek v4 Flash benchmark compares coding performance against Qwen3.6-27B, 3.5-122B, and Gemma 4 31B using a two-shot file-editing task subset
r/LocalLLaMA
→ story
Aug 4, 2026
Gemma 4 on 500MB
r/LocalLLaMA
→ story
Aug 4, 2026
Cost-effective automated judging of natural-language mathematical proofs using open-weight models aligns with human grading decisions on IMO-GradingBench instances
arXiv cs.CL
→ story
Aug 3
Aug 3, 2026
KAT Coder 2.5 dev: claims faster and more accurate than Qwen 3.6 35b; benchmarks and user impressions cited
r/LocalLLaMA
→ story
Aug 3, 2026
Knowledge distillation in small instruction-tuned LMs has asymmetric effects on bias, improving context-following for unambiguous tasks but degrading calibration for ambiguous tasks.
arXiv cs.CL
→ story
Aug 2
Aug 2, 2026
Parlor v2: best-effort fully local GPT-Live clone on an M3 Pro
r/LocalLLaMA
→ story
Aug 1
Aug 1, 2026
Tomte harness for Gemma 4 released; enables fast local inference on Macs with M processors
r/LocalLLaMA
→ story
←
1
2
→