LIVE · refreshes every 20 min
updated Aug 12, 02:02 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
110010
.art
/ topic
CUDA
Aug 10
1d ago
Ante 0.2 enables a self-contained coding agent that manages llama.cpp offline by pointing at a GGUF; it handles the inference engine, supported on Metal, CUDA, Vulkan, or CPU, with pinned builds and model discovery.
r/LocalLLaMA
→ story
2d ago
Retired engineer develops a ML programming language to teach and demonstrate ML concepts with in-browser training and CUDA/NVIDIA and MLX/Apple Silicon support
r/LocalLLaMA
→ story
Aug 8
3d ago
Building a zero-dependency C inference engine for BitNet (1.58-bit); achieves 36 tok/s on a Xeon CPU
r/LocalLLaMA
→ story
3d ago
Qwen3.5-9B MTP GGUF shows 35.8% acceptance on CUDA and 91–92% on Vulkan; two Radeon/RADV systems are 77% and 128% faster on Vulkan.
r/LocalLLaMA
→ story
Aug 7
4d ago
GPU upgrade plans for RTX 3090 to two ASRock AMD Pro R9700s face power-connector challenges; author questions viability of 16-pin adapter solution.
r/LocalLLaMA
→ story
Aug 5
7d ago
PSA: Update CUDA from 13.2 to 13.3 to fix DeepSeek V4 Flash 0731 looping issue
r/LocalLLaMA
→ story
Aug 4
7d ago
Full 1M context on a single RTX5090 with DDR5 desktop setup using vLLM CPU/RAM offloading; ~800 tps per prompt and 15+ tps decode
r/LocalLLaMA
→ story
Aug 3
8d ago
Nvidia's CUDA faces new threats from AI coding agents.
Hacker News (AI)
→ story
Jul 31
11d ago
AMD advances AI in 2026; can it break the CUDA moat?
Hacker News (AI)
→ story
11d ago
Apache-2.0 Tritium releases an open source ternary LLM engine in Rust/CUDA for quantization, serving, and training on consumer GPUs
r/LocalLLaMA
→ story
Jul 30
12d ago
Inkling-Small by thinkingmachines: 276B total parameters, 12B active, 1M context window.
r/LocalLLaMA
→ story
12d ago
Does MTP head load in VRAM by default?
r/LocalLLaMA
→ story
12d ago
2× Radeon AI PRO R9700 GPUs chosen for a local AI server; user questions whether AMD/ROCm is a realistic choice today compared to NVIDIA/CUDA for running local LLMs
r/LocalLLaMA
→ story
Jul 29
13d ago
K-Search translates CUDA optimization knowledge into architecture-native MLX strategies for Apple Silicon
BAIR (Berkeley)
→ story
13d ago
Kernel Forge: an agent harness for LLM-based generation and optimization of CUDA kernels
arXiv cs.AI
→ story
Jul 28
14d ago
Llama.cpp updates include chunked SSD matmul acceleration for Mamba-2 prefill and a ggml-metal FWHT kernel for the metal backend
r/LocalLLaMA
→ story
Jul 27
15d ago
Parameter-free adaptive sparse attention via compression-based content selection
arXiv cs.LG
→ story
Jul 26
16d ago
POCKET-35B agentic model runs on CPU at 59 t/s using GGUF on stock llama.cpp
r/LocalLLaMA
→ story
Jul 25
17d ago
Benchmarks: TensorSharp vs. llama.cpp
r/LocalLLaMA
→ story
Jul 24
18d ago
SonicSampler: unified tile-aware kernels for LLM sampling and speculative verification
arXiv cs.AI
→ story
Jul 21
21d ago
KernelBench-Verified: LLM-generated kernels often inflate performance; evaluation frameworks must evolve to measure true speedup, not reported gains.
arXiv cs.LG
→ story