LIVE · refreshes every 20 min
updated Sep 4, 16:22 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
110010
.art
/ topic
CUDA
Sep 2
1d ago
The Modern CUDA Toolbox in Practice: A Step-by-Step Optimization Walkthrough
NVIDIA Developer
→ story
2d ago
CUDA-Harness enables agentic CUDA kernel generation and optimization from natural language
arXiv cs.CL
→ story
Aug 18
16d ago
Bonsai and Ternary Bonsai updates: 1-bit, 1.58-bit (Ternary), and 27B models; CUDA and Vulkan backends merged on mainline with optimization PRs improving throughput by 15–40% and perf by 8%
r/LocalLLaMA
→ story
Aug 16
18d ago
Llama.cpp upgrades ROCm to 7.14 for Radeon 780m iGPU benchmarks; ROCm 7.14 adds gfx1103 support, but prebuilt binaries lack the target, requiring a source build
r/LocalLLaMA
→ story
18d ago
Qwen3.8 27B Q2 vs Q3 and Qwen3.6 35B-A3B MoE on 12GB VRAM
r/LocalLLaMA
→ story
Aug 10
25d ago
Ante 0.2 enables a self-contained coding agent that manages llama.cpp offline by pointing at a GGUF; it handles the inference engine, supported on Metal, CUDA, Vulkan, or CPU, with pinned builds and model discovery.
r/LocalLLaMA
→ story
25d ago
Retired engineer develops a ML programming language to teach and demonstrate ML concepts with in-browser training and CUDA/NVIDIA and MLX/Apple Silicon support
r/LocalLLaMA
→ story
Aug 8
26d ago
Building a zero-dependency C inference engine for BitNet (1.58-bit); achieves 36 tok/s on a Xeon CPU
r/LocalLLaMA
→ story
27d ago
Qwen3.5-9B MTP GGUF shows 35.8% acceptance on CUDA and 91–92% on Vulkan; two Radeon/RADV systems are 77% and 128% faster on Vulkan.
r/LocalLLaMA
→ story
Aug 7
28d ago
GPU upgrade plans for RTX 3090 to two ASRock AMD Pro R9700s face power-connector challenges; author questions viability of 16-pin adapter solution.
r/LocalLLaMA
→ story
Aug 5
Aug 5, 2026
PSA: Update CUDA from 13.2 to 13.3 to fix DeepSeek V4 Flash 0731 looping issue
r/LocalLLaMA
→ story
Aug 4
Aug 4, 2026
Full 1M context on a single RTX5090 with DDR5 desktop setup using vLLM CPU/RAM offloading; ~800 tps per prompt and 15+ tps decode
r/LocalLLaMA
→ story
Aug 3
Aug 3, 2026
Nvidia's CUDA faces new threats from AI coding agents.
Hacker News (AI)
→ story
Jul 31
Jul 31, 2026
AMD advances AI in 2026; can it break the CUDA moat?
Hacker News (AI)
→ story
Jul 31, 2026
Apache-2.0 Tritium releases an open source ternary LLM engine in Rust/CUDA for quantization, serving, and training on consumer GPUs
r/LocalLLaMA
→ story
Jul 30
Jul 30, 2026
Inkling-Small by thinkingmachines: 276B total parameters, 12B active, 1M context window.
r/LocalLLaMA
→ story
Jul 30, 2026
Does MTP head load in VRAM by default?
r/LocalLLaMA
→ story
Jul 30, 2026
2× Radeon AI PRO R9700 GPUs chosen for a local AI server; user questions whether AMD/ROCm is a realistic choice today compared to NVIDIA/CUDA for running local LLMs
r/LocalLLaMA
→ story
Jul 29
Jul 29, 2026
K-Search translates CUDA optimization knowledge into architecture-native MLX strategies for Apple Silicon
BAIR (Berkeley)
→ story
Jul 29, 2026
Kernel Forge: an agent harness for LLM-based generation and optimization of CUDA kernels
arXiv cs.AI
→ story
Jul 28
Jul 28, 2026
Llama.cpp updates include chunked SSD matmul acceleration for Mamba-2 prefill and a ggml-metal FWHT kernel for the metal backend
r/LocalLLaMA
→ story
Jul 27
Jul 27, 2026
Parameter-free adaptive sparse attention via compression-based content selection
arXiv cs.LG
→ story
Jul 26
Jul 26, 2026
POCKET-35B agentic model runs on CPU at 59 t/s using GGUF on stock llama.cpp
r/LocalLLaMA
→ story
Jul 25
Jul 25, 2026
Benchmarks: TensorSharp vs. llama.cpp
r/LocalLLaMA
→ story
Jul 24
Jul 24, 2026
SonicSampler: unified tile-aware kernels for LLM sampling and speculative verification
arXiv cs.AI
→ story
Jul 21
Jul 21, 2026
KernelBench-Verified: LLM-generated kernels often inflate performance; evaluation frameworks must evolve to measure true speedup, not reported gains.
arXiv cs.LG
→ story