LIVE · refreshes every 20 min
updated Aug 23, 12:41 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
110010
.art
/ archive
Research
Aug 12
11d ago
Agent-MD: Selective LLM intervention with event-driven escalation for stateful GCMC–MD campaigns
arXiv cs.AI
→ story
11d ago
AI scientist for studying generalization in quadruped robot navigation policies in simulation; addresses drift in autonomous research loops
arXiv cs.AI
→ story
11d ago
Analyzing LLM alignment across single- and multi-cultural settings using cultural consensus theory
arXiv cs.CL
→ story
11d ago
ASR-roundtrip evaluation can mask context- and convention-dependent reading errors in Chinese news TTS
arXiv cs.CL
→ story
11d ago
Asymmetric framing of La France insoumise and Rassemblement National in French news headlines, 2022–2025
arXiv cs.CL
→ story
11d ago
Automating and scaling behavioral scientific research on AI agents with AEROBAT
arXiv cs.AI
→ story
11d ago
Beyond detection: evaluating defensive LLMs against AI-generated social engineering in live turn-by-turn interaction
arXiv cs.AI
→ story
11d ago
Boundary-Seeking Policy Gradient for safe reinforcement learning
arXiv cs.LG
→ story
11d ago
Calibrating post-training feature shifts for LLM data contamination detection
arXiv cs.CL
→ story
11d ago
CHORUS: complementary experts for high-coverage testbench stimulus generation
arXiv cs.AI
→ story
11d ago
ChronoSSM: training for temporally aware representations in autoregressive state space models
arXiv cs.LG
→ story
11d ago
Closed-loop LLM co-pilots for digital agriculture
arXiv cs.AI
→ story
11d ago
CMU-Drive and V2V-VLA: cooperative multi-agent unified driving with reasoning benchmark and vehicle-to-vehicle vision-language-action models
arXiv cs.AI
→ story
11d ago
Contextual value alignment via multilayer combinatorial fusion
arXiv cs.AI
→ story
Aug 11
11d ago
CARE-X aims for clinically useful radiology VLMs with auxiliary supervision, reward-aligned learning, and tool-augmented measurement
Microsoft Research
→ story
12d ago
LUCID: Latent-Skill Unified Control via Imagined Dynamics for long-horizon humanoid loco-manipulation
arXiv cs.LG
→ story
12d ago
Many Are My Names": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders
arXiv cs.CL
→ story
12d ago
Multilingual supervision improves figurative language identification in proverbs, evaluated across seven languages and 742 proverb concepts
arXiv cs.CL
→ story
12d ago
NeuPAT: neuron-aware plasticity allocation tuning for language-preserving mllms
arXiv cs.CL
→ story
12d ago
Neural operators for immersed-boundary soft swimmers locomotion
arXiv cs.LG
→ story
12d ago
PhysAttNet: enhancing predictive performance in industrial and astrophysical time series via physics-informed attention
arXiv cs.LG
→ story
12d ago
Policy learning with mu-resets: the sample complexity under policy realizability for the Kakade–Langford interaction protocol
arXiv cs.LG
→ story
12d ago
Polish vision-language evaluation PoVisLE assesses cultural grounding in vision-language models
arXiv cs.CL
→ story
12d ago
Prompt Embedding Probes detects hallucinations from hidden states of a frozen LLM
arXiv cs.CL
→ story
12d ago
Reasoning models fail to ration test-time compute across questions
arXiv cs.CL
→ story
12d ago
Scaling inherently interpretable language models by training with interpretability as a constraint alongside language modeling objective
arXiv cs.CL
→ story
12d ago
Shape mutating expert compression: LorExperts and BTExperts
arXiv cs.LG
→ story
12d ago
SkillConsist detects inconsistencies in agent skills via bidirectional graph alignment
arXiv cs.LG
→ story
12d ago
SPECTRA: pushing the KV cache beyond the 2-bit cliff via spectral transform coding
arXiv cs.LG
→ story
12d ago
Spectral outliers reveal dominant learned structure in transformer attention.
arXiv cs.LG
→ story
12d ago
STEMMA: an adversarial multi-agent framework for evaluating self-identity consistency in LLMs
arXiv cs.CL
→ story
12d ago
SurakshaEval: a safety benchmark for multilingual LLMs covering ten major Indian languages
arXiv cs.CL
→ story
12d ago
SurveyReview: a reviewer-aligned benchmark for survey evaluators
arXiv cs.CL
→ story
12d ago
TEMPER: Tensorized Efficient Manifold-constrained Parameterization for Expressive Residual Routing
arXiv cs.LG
→ story
12d ago
The no-meaning falsity: the structural impossibility of the arbitrary sign in Classical Arabic
arXiv cs.CL
→ story
12d ago
Trace-driven evaluation can mislead MoE expert caching due to replay semantics, workload contamination, and operating regimes.
arXiv cs.LG
→ story
12d ago
Tracing sources of epistemic uncertainty in deep learning predictions with homo- and heteroscedastic linearized estimators
arXiv cs.LG
→ story
12d ago
Unified hallucination fuzzing for multimodal large language models
arXiv cs.CL
→ story
12d ago
V-Simba improves architectural efficiency for reinforcement learning in visual continuous control
arXiv cs.LG
→ story
12d ago
WuYuEval: a multi-level benchmark for evaluating large language models in solid waste management across foundational knowledge, domain reasoning, and expert decision-making
arXiv cs.CL
→ story
←
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
→