All Articles
16096 articles total
HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models
Introduces HiFloat4 and Rollout-ResQ to enable end-to-end 4-bit precision RL post-training for LLMs with minimal accuracy loss.
A Graph-Native Bitemporal Memory Store for Conversational AI Agents
A bitemporal memory store for AI agents using Neo4j property graphs and HNSW vector indexes to support point-in-time semantic retrieval.
Danube's record low levels force shutdown of Hungary's only nuclear plant
Hungary's only nuclear power plant has been forced to shut down due to record low water levels in the Danube river.
Show HN: A udev implementation in Guile Scheme
A developer has implemented udev, the Linux device manager, using the Guile Scheme programming language.
DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
An analysis of the intelligence, performance, and pricing of the DeepSeek V4 Flash 0731 model.
Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
Proposed game-theoretic framework for fine-tuning LLMs that optimizes the trade-off between reward maximization and statistical detectability from a reference policy.
Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
Research on classifying inconsistencies in financial disclosure texts, comparing encoders and LLMs for fine-grained conflict detection.
Zero-Fi: Zero-Shot Wi-Fi-Based Human Activity Recognition via Contrastive Signal-Language Alignment
Zero-Fi is a new framework for zero-shot human activity recognition using Wi-Fi signals aligned with natural language descriptions.
Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction
A study showing that certain algorithmic collusion strategies in pricing are undetectable by current auditing methods, even in LLM-based agents.
Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
Research suggesting that LLM misalignment during fine-tuning manifests as a shift in 'personality' traits based on the Big Five model.
Voice Memory for Agentic Speech Recognition
Voice Memory introduces a listener-thinker architecture for speech recognition that uses an auditable memory file instead of weight updates for correction.
FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing
FAS-R1 is a multi-task MLLM for face anti-spoofing that combines a long-CoT dataset with GRPO post-training for better reasoning.
Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Google reports a significant increase in Chrome bug fixes in June, attributing the surge to the use of AI.
Entity Resolution in Practice: Lessons from a Self-Serve Pipeline
Researchers share practical lessons from building a self-serve entity resolution system, emphasizing the need for multiple algorithms and separate precision/recall fixes.
AgentGUI: An Interface for Observing and Steering Long-Running AI Agents
AgentGUI is introduced as a locally hosted interface for observing and steering long-running AI agents with rich trajectory visualizations.
SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation
The SARC-DQ framework addresses data-quality gating for agentic AI to prevent costly actions based on stale or incorrect metadata.
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
StealthBench is a new benchmark for measuring operational stealth and OPSEC tradecraft in autonomous offensive-security agents.
Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
Cognitive Diagnostic Profiling (CDP) is proposed to align LLM-simulated examinees with real human psychometric data for better test calibration.
Automorphism-Induced Non-Canonicity in Top-k Explanations of Graph Neural Networks
Researchers identify a structural obstruction in Graph Neural Network explainers where automorphisms lead to arbitrary top-k edge reports.
When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
A cross-domain benchmark reveals that LLM-simulated synthetic users fail to accurately replicate human survey responses, especially across cultures.