All Articles
17659 articles total
SoK: Attack and Defense Landscape of Mobile On-device AI Systems
A comprehensive systematization of knowledge regarding the attack and defense landscape of mobile on-device AI (MoAI) systems.
Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis
A new unified guidance framework for Flow Matching (FM) that reduces inference latency and improves speaker similarity in speech synthesis.
When AI meets quantum information: A comprehensive review
A review paper exploring the bidirectional relationship between artificial intelligence and quantum information, discussing AI's role in designing quantum systems and QI's impact on AI models.
MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts
Introduction of MEPA, a scale-aware token-routed Mixture of Experts (MoE) architecture for improving Visual AutoRegressive modeling in image generation.
Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
Presentation of FoCo, a framework for Zero-Shot Composed Image Retrieval that uses proxy tasks to better model semantic modifications.
MalariAI: A Label-Resilient Decoupled Framework for Universal Cell Segmentation and Explainable Stage Classification in Dense Malaria Blood Smears
MalariAI is a decoupled framework for cell segmentation and stage classification in malaria blood smears, designed for use in resource-limited settings.
Learning Generalizable Skill Policy with Data-Efficient Unsupervised RL
GenDa is a unified framework for unsupervised reinforcement learning that improves data efficiency and generalizability through skill relabeling and an Information Bottleneck.
NeuroCogMap Reveals Cognitive Organization of Large Language Models
NeuroCogMap is a framework that maps the internal functional organization of LLMs to cognitive capabilities and relates them to human cortical responses.
Aerial Photographs (2017)
A link to a collection of aerial photographs from 2017.
Z.ai launches ZCode to challenge Cursor, Claude Code and GitHub Copilot in AI coding
Z.ai launches ZCode, an agentic development environment powered by the open-source GLM-5.2 model trained on Huawei silicon.
Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds
Researchers introduce a four-stage diagnostic to test LLM physics literacy in parallel, counterfactual physical worlds to distinguish reasoning from recall.
What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models
A new framework using Planning KL-divergence (PKL) allows VLMs to identify and reason about planning-critical occluded agents in autonomous driving.
An LLM-Based Framework for Intent-Driven Network Topology Design
This research explores an LLM-based framework for generating resilient network topologies from natural language requirements using a constraint-driven pipeline.
Learning When to Listen: Gated Affect Fusion for Human Motion Prediction
The Gated Affect Transformer (GAT) is introduced to dynamically regulate facial affect information for improved human motion prediction in real-world videos.
Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions
An empirical survey analyzing the bias-reliability tradeoff in LLM evaluation systems across eleven different evaluator-agent conditions.
RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail
The RetailSMV dataset is introduced to study the adaptation of foundation video world models in retail environments using egocentric and exocentric views.
K-Inverse-RFM: A Modified RFM that Bridges the Gap to Neural Networks for Data-Corrupted Mathematical Tasks
K-Inverse-RFM is proposed as a modified Recursive Feature Machine that closes the performance gap with neural networks on data-corrupted mathematical tasks.
DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning
DiscoLoop is a new looping architecture that uses both discrete embedding and continuous hidden-state channels to improve multi-hop reasoning in LLMs.
Amazon has enough satellites to launch its Starlink competitor
Amazon Leo has deployed 396 satellites, putting it on track for commercial availability by mid-2026 to compete with SpaceX's Starlink.
EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards
Introduction of EgoSafetyBench, a diagnostic egocentric video benchmark designed to evaluate Vision-Language Models (VLMs) as runtime safety guards for embodied AI.