LIVE · refreshes every 20 min
updated Aug 26, 19:42 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
110010
.art
/ topic
Llama
Jul 27
Jul 27, 2026
Tool llmux automates parameter management for vLLM and llama.cpp model swaps by saving per-model profiles and launching the appropriate engine version
r/LocalLLaMA
→ story
Jul 27, 2026
Llama.cpp signals a forthcoming change requiring regeneration of all GGUFs generated before the update
r/LocalLLaMA
→ story
Jul 26
Jul 26, 2026
Minimax M3 support with MSA has been merged into llama.cpp
r/LocalLLaMA
→ story
Jul 26, 2026
BeeLlama.cpp v0.4.1 adds KVarN, KV cache precision tail, and q2_0-q3_1/q6_0 cache support with improved benchmarks and VRAM efficiency
r/LocalLLaMA
→ story
Jul 26, 2026
90 agentic bakeoff compares ThinkingCap, Fable Fusion, and stock Qwen3.6-27B across 90 runs with 6 self-grading tasks and 5 reps per model
r/LocalLLaMA
→ story
Jul 26, 2026
POCKET-35B agentic model runs on CPU at 59 t/s using GGUF on stock llama.cpp
r/LocalLLaMA
→ story
Jul 26, 2026
GLM 5.2 on a 4-socket Xeon box with 1TB RAM and RTX 3060; ik_llama.cpp runs on CPU with 24 attention layers on GPU, but generation crashes with NaN logits beyond 32–64k context
r/LocalLLaMA
→ story
Jul 25
Jul 25, 2026
Llama.cpp now has full MCP support.
r/LocalLLaMA
→ story
Jul 25, 2026
Choosing a model for Hermes on a 4x 3090, 96 GB VRAM setup; llama.cpp or vllm, for a personal AI playground with a few family users
r/LocalLLaMA
→ story
Jul 25, 2026
Benchmarks: TensorSharp vs. llama.cpp
r/LocalLLaMA
→ story
Jul 22
Jul 22, 2026
Interpreting how instruction-tuned transformers encode discourse relations, focusing on causation and antithesis, in next-token prediction tasks
arXiv cs.CL
→ story
Jul 21
Jul 21, 2026
High-accuracy low-bit KV-cache quantization via local distribution restoration
arXiv cs.LG
→ story
Jul 20
Jul 20, 2026
Show HN: A fast, free AI text humanizer powered by Groq Llama 3.3
Hacker News (AI)
→ story
Jun 17
Jun 17, 2025
Announcing the inaugural Llama Startup Program cohort - AI at Meta
Meta AI (via Google News)
→ story
Apr 5
Apr 5, 2025
Llama 4 marks a new era of natively multimodal AI innovation at Meta
Meta AI (via Google News)
→ story
Nov 20
Nov 20, 2024
Meta builds Llama-based chatbot to engage stakeholders of an intergovernmental organization
Meta AI (via Google News)
→ story
Jul 23
Jul 23, 2024
Meta introduces Llama 3.1, its most capable models to date
Meta AI (via Google News)
→ story
Apr 18
Apr 18, 2024
Meta Llama 3 advances as the most capable openly available LLM to date
Meta AI (via Google News)
→ story
Jul 18
Jul 18, 2023
Meta and Microsoft introduce the next generation of Llama; AI at Meta
Meta AI (via Google News)
→ story
←
1
2
3
4
→