All Articles
17919 articles total
How to invest when everything is moving too fast
A discussion between AI investors regarding investment strategies in a rapidly evolving market.
PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models
Introduction of PHANTOM, a large-scale open-source dataset of multimodal adversarial attacks designed to test the robustness of Vision-Language Models (VLMs).
Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War
Age of LLM is a 1v1 strategic benchmark for LLMs featuring a grid-based game with fog of war, diplomacy, and strict schema reliability requirements.
ATRIA: Adaptive Traceable ECG Reporting with Iterative Agents
ATRIA is a multi-agent ECG reporting system that utilizes an iterative workflow to allow clinicians to verify and revise individual findings.
Cycle-Consistent Neural Explanation of Formal Verification Certificates
Researchers propose a cycle-consistent neural architecture to translate formal verification certificates into natural language, achieving higher soundness and faster inference than general LLMs.
Agentic AI for Bilevel Long-Term Optimization of Policy-Driven Physical Layer Systems
Agentic-LTPO is a nested bilevel optimization framework using agentic AI to optimize long-term performance in physical layer systems like cell-free MIMO beamforming.
Can Aggregate Invariants Accelerate Continuous Subgraph Matching? Limits, Laws, and a Dynamic Spectral Index
The authors study the use of aggregate invariants and spectral filtering to accelerate continuous subgraph matching in dynamic graphs and release a dynamic local-spectra index.
ReM-MoA: Reasoning Memory Sustains Mixture-of-Agents Scaling
ReM-MoA is a memory-augmented Mixture-of-Agents framework that uses Ranked Reasoning Memory to sustain scaling gains as pipeline depth increases.
Bayesian control for coding agents
The paper introduces a Bayesian controller for coding agents to optimize the decision of when to refine, verify, or stop based on cost-sensitive sequential hypothesis testing.
"Fix" MacBook Neo Cursor Lag: Record 1 Pixel of the Screen Every 10 Seconds
A discussion about a peculiar workaround for cursor lag on MacBook Neo, involving recording a small portion of the screen every few seconds.
Remaking BBC test cards to teach you video processing
An educational project focused on recreating BBC test cards to teach the fundamentals of video processing.
Towards Federated Long-Tailed Graph Learning: An Energy-Guided Dual Decoupling Approach
Introduction of FedEPD, a federated graph learning framework designed to handle long-tailed data distributions via dual decoupling of topological purification and semantic recalibration.
Probing the Misaligned Thinking Process of Language Models
Research on detecting misaligned behaviors in LLMs (e.g., strategic deception) by using linear probes on internal activations to identify cognitive indicators.
Tractable Reasoning and Conjunctive Query Answering for Defeasible DL-Lite under Rational Closure
A study providing a plug-in architecture for efficient reasoning and conjunctive query answering in the DL-Lite family of description logics under Rational Closure.
LemonHarness Technical Report
Presentation of LemonHarness, an execution framework for long-horizon LLM agents that uses explicit workspace boundaries and time-aware execution to improve stability.
Prob-BBDM: a Probabilistic Brownian Bridge Diffusion Model for MRI sequence image-to-image translation
Prob-BBDM is a probabilistic Brownian Bridge Diffusion Model for efficient 2D-to-3D MRI sequence image-to-image translation in medical imaging.
MVG-KAN: Multi-View Geo-Wind Guided KAN for PM$_{2.5}$ Forecasting
MVG-KAN is a multi-view model using Geo-Wind graphs and Kolmogorov-Arnold Networks (KAN) for accurate short-term PM2.5 air quality forecasting.
Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation
Introduction of DigenRL, a disaggregated RL framework for diffusion-based generative LLMs that increases throughput via pipeline and time-step parallelism.
When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs
Research showing that LLMs suppress 'Causal Caution' in practical advisory contexts compared to academic ones, suggesting a need for separate causal auditing agents.
VeryTrace: Verifying Reasoning Traces through Compilable Formalism and Structured Verification
VeryTrace is a zero-shot framework that uses a Domain-Specific Language (DSL) to formalize and verify the reasoning steps of LLMs to reduce hallucinations.