LIVE · refreshes every 20 min
updated Sep 4, 01:23 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
110010
.art
/ topic
DeepSeek
Aug 9
25d ago
endless-frontier/BigBang-v1 - qwen 3.5 finetunes benchmark results; comparisons between DeepSeek Flash and Pro show per-benchmark variance
r/LocalLLaMA
→ story
25d ago
DeepSeek-V4-Flash-0731-Latent-Reasoning presents a model thinking in latent space
Hacker News (AI)
→ story
25d ago
DeepSeek V4 Flash 0731 achieves 82.7% on Terminal-Bench 2.1 in independent public-harness run (445 trials)
r/LocalLLaMA
→ story
25d ago
Memory bandwidth limits observed on Intel Sapphire Rapids with DDR5-4800 RDIMMs; user reports ~36–40 GB/s vs theoretical 153 GB/s for DeepSeek-V4-Flash-0731 MoE workload
r/LocalLLaMA
→ story
Aug 8
26d ago
DSpark draft model runs extremely slowly (1-2 t/s) with DeepSeek-V4-Flash on llama-server vs MTP
r/LocalLLaMA
→ story
26d ago
Quants of deepseek flash 0731: recommendations for fast public benchmarks to test quantization effects with about 1 million tokens
r/LocalLLaMA
→ story
26d ago
DSV4F 0731 praised as a reliable workhorse with strong benchmarks for Hermes agent tasks, OpenCode coding, and long-running integrations
r/LocalLLaMA
→ story
26d ago
DeepSeek-V4-Flash shown unreliable for non-coding tasks, user reports; performance unreliable beyond coding
r/LocalLLaMA
→ story
Aug 7
27d ago
Serving Deepseek v4 Flash 0731 on 2x DGX Spark; how to lower VRAM usage and increase OS headroom for RAM?
r/LocalLLaMA
→ story
27d ago
DeepSeek V4 Flash 0731 reveals ARC-AGI results
r/LocalLLaMA
→ story
27d ago
Is GLM 5.2 and Kimi 2.7 still relevant amid newer models like Kimi K3, Qwen 3.8 Max, and V4 pro Deepseek?
r/LocalLLaMA
→ story
Aug 6
28d ago
Deepseek aims to run LLMs locally without a GPU; user argues it as an affordable alternative
r/LocalLLaMA
→ story
28d ago
Frontier narrows pricing gap as performance catches up to price concerns; users report downgraded free version after new models release
r/LocalLLaMA
→ story
28d ago
DeepSeek signals a significant price increase for AI services, testing its low-cost edge
Hacker News (AI)
→ story
28d ago
AI labs raise prices with little notice after DeepSeek API hike; customers warn about production impact and the need to reconsider hosting or providers.
r/LocalLLaMA
→ story
28d ago
Users discuss training their own AI from scratch on personal systems for experimentation, sharing hardware and techniques tested from recent research.
r/LocalLLaMA
→ story
Aug 5
29d ago
I remember a time when 'flash' meant 32B
r/LocalLLaMA
→ story
29d ago
NVIDIA V100 32GB suitable for DeepSeek V4 Flash for 30-50 users, per user inquiry
r/LocalLLaMA
→ story
29d ago
DeepSeek's Liang Wenfeng: Full remarks from an investor meeting
Hacker News (AI)
→ story
29d ago
DeepSeek-v4-flash-mini demonstrates continued exploration of whoops DeepSeek work; user submission discusses extending DeepSeek via exploration
r/LocalLLaMA
→ story
29d ago
TensorSharp adds MoE CPU-offload feature to main, enabling up to 35B-A3B MoE to fit with long-context KV cache on 12-16 GB GPUs; router, shared experts stay on accelerator.
r/LocalLLaMA
→ story
Aug 5, 2026
HuggingFace model requires more DRAM; user anticipates DRAM prices may rise
r/LocalLLaMA
→ story
Aug 5, 2026
PSA: Update CUDA from 13.2 to 13.3 to fix DeepSeek V4 Flash 0731 looping issue
r/LocalLLaMA
→ story
Aug 4
Aug 4, 2026
DeepSeek v4 Flash benchmark compares coding performance against Qwen3.6-27B, 3.5-122B, and Gemma 4 31B using a two-shot file-editing task subset
r/LocalLLaMA
→ story
Aug 4, 2026
Full 1M context on a single RTX5090 with DDR5 desktop setup using vLLM CPU/RAM offloading; ~800 tps per prompt and 15+ tps decode
r/LocalLLaMA
→ story
Aug 4, 2026
DeepSeek V4 Flash 2.98x faster, lossless; updated template supports reasoning levels
Hacker News (AI)
→ story
Aug 4, 2026
Best way to run DS4 flash on a Mac (M3 Ultra) with 192 GB+ VRAM; includes dspark/mtp support and rising token throughput
r/LocalLLaMA
→ story
Aug 4, 2026
Cost-effective automated judging of natural-language mathematical proofs using open-weight models aligns with human grading decisions on IMO-GradingBench instances
arXiv cs.CL
→ story
Aug 3
Aug 3, 2026
Ling-3.0-flash tested as an additional model option before qwen3.8 27b
r/LocalLLaMA
→ story
Aug 3, 2026
Was the release of deepseek v4 flash planned to take spotlight against 5.6 luna?
r/LocalLLaMA
→ story
Aug 3, 2026
Dual RTX Pro 6000s: DSpark with SGLang and VLLM won’t work; user seeks configuration tips and recipes to get it running
r/LocalLLaMA
→ story
Aug 3, 2026
Question about quantization versus model size in LLMs
r/LocalLLaMA
→ story
Aug 2
Aug 2, 2026
Quantizing KV Cache for DeepSeek V4 Flash degrades quality, per DS4F perplexity results
r/LocalLLaMA
→ story
Aug 2, 2026
DeepSeek-V4-Flash-0731 surpasses Fable-5, Sol, and Kimi-K3 on Chess Benchmark
r/LocalLLaMA
→ story
Aug 2, 2026
Setting up a 16x GX10 (DGX Spark) cluster to run locally frontier-level open models with 8x per 2-model split and potential 2T+ models
r/LocalLLaMA
→ story
Aug 2, 2026
DSv4F prompts should not include system messages mid-conversation to avoid corrupting the prompt cache
r/LocalLLaMA
→ story
Aug 2, 2026
DeepSeek's Theory of the AI Gap
Hacker News (AI)
→ story
Aug 2, 2026
June newsletter: highlights include accidental cyberattacks by OpenAI and Anthropic models under test, GPT-5.6 Sol/Terra/Luna, Claude Opus 5, Kimi K3 and DeepSeek-V4-Flash-0731, and other model releases
Simon Willison
→ story
Aug 1
Aug 1, 2026
Can smaller models keep getting smarter, or is there a limit to reducing size without losing intelligence?
r/LocalLLaMA
→ story
Aug 1, 2026
Deepseek v4 flash 0731 struggles to follow rules prompts and skills
r/LocalLLaMA
→ story
←
1
2
3
4
→