All Articles
17859 articles total
HauntAttack: When Attack Follows Reasoning as a Shadow
Research introducing HauntAttack, a black-box adversarial attack framework targeting the internal reasoning processes of Large Reasoning Models.
DMSC: Dynamic Multi-Scale Coordination Framework for Time Series Forecasting
DMSC, a dynamic multi-scale coordination framework for time series forecasting that utilizes MoE and adaptive patch decomposition.
The US lifts its block on Mythos 5
The US government has lifted its block on Mythos 5.
My favorite Govee smart lamps are at their lowest prices ever for Prime Day
The Verge reports on Prime Day deals for Govee smart lamps, highlighting features like Matter support and music sync.
South Korea plans to train entire military as "drone warriors"
South Korea plans to train its entire military force as 'drone warriors,' treating drones as a universal combat tool.
New agentic memory framework uses 118K tokens per query. LangMem burns through 3.26M.
Researchers developed MRAgent, a framework that uses active memory reconstruction to significantly reduce token consumption and runtime costs for AI agents.
Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
This paper identifies 'library drift' in self-evolving LLM skill libraries and proposes a governance recipe to fix retrieval degradation and performance stagnation.
Augmentation techniques for video surveillance in the visible and thermal spectral range
A study investigates the use of multispectral CNN-based object detection and augmentation techniques for video surveillance using visible and thermal cameras.
Automated reproducibility assessments in the social and behavioral sciences using large language models
Researchers demonstrate that LLMs can be used as scalable screening tools to automate reproducibility assessments in social and behavioral sciences.
A-Evolve-Training: Autonomous Post-Training of a 30B Model
NVIDIA researchers report A-Evolve-Training, an autonomous post-training system for 30B-550B models that can recursively self-improve without human intervention.
Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning
The 'Cliff Tokens' research identifies single-token failure triggers in LLM mathematical reasoning and proposes Cliff-DPO to improve accuracy.
Autodata: An agentic data scientist to create high quality synthetic data
Autodata introduces an agentic data scientist method to create high-quality synthetic training and evaluation data through meta-optimization.
The open source DOCX editor submitted to HN a few weeks ago has been deleted
A previously submitted open source DOCX editor has been deleted from Hacker News.
Lippmann Photography
A post regarding Lippmann Photography, likely unrelated to technical development.
A Concept of Possibility for Real-World Events
A research paper proposing a new concept of 'possibility' for real-world events to improve planning and feasibility analysis in AI.
Human-AI Complementarity: A Goal for Amplified Oversight
Research on 'Amplified Oversight', exploring how AI assistants can improve human fact-verification of AI outputs without causing over-reliance.
SciFig: Towards Automating Editable Figure Generation for Scientific Papers
Introduction of SciFig, a multi-agent framework for generating editable methodology figures for scientific papers, including a new benchmark and evaluation protocol.
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
Proposed Joint Reward Modeling (JRM) which optimizes preference learning and language modeling together to create efficient and accurate visual reward models.
CLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts
HIPE-2026, a CLEF evaluation lab focused on extracting person-place relations from multilingual historical texts for knowledge-graph construction.
Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents
A unifying framework called the Experience Compression Spectrum that categorizes agent memory, skills, and rules by their level of compression.