All Articles
17521 articles total
ORCAID: Oblique Rule-Based Continuous-Action Interpretation for Deep RL Policies
The researchers introduce ORCAID, a method for extracting interpretable rule-based policies from deep RL agents in mixed continuous-discrete environments.
FMMVCC: Fuzzy Mamba-based Multi-View Contrastive Clustering for Univariate Time Series
FMMVCC is a new Mamba-based deep clustering framework for univariate time series that uses state space sequence modeling for efficient temporal representation.
Bayesian Optimization of Genetic Algorithm Hyperparameters in a Multi-Fidelity Framework for Efficient Lattice Material Design
A multi-fidelity framework combining Bayesian Optimization and Genetic Algorithms is proposed for the efficient design of lattice materials.
CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street-View Semantics and Temporal Context Training
CarbonCLIP is a multimodal distillation framework that enhances satellite-based carbon emission prediction by integrating street-view semantics and temporal context.
Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders
The MM-VAP framework improves turn-taking prediction for social robots by projecting future voice activity using multimodal audio-visual inputs.
POO-LPSP: Parallel Osprey Optimized Least Penalty-Squared Prioritization Methods for Priority Derivation in the Analytic Hierarchy Process
The POO-LPSP method utilizes a Parallel Osprey Optimization Algorithm to solve complex prioritization models within the Analytic Hierarchy Process.
FedCVESA: Taking Away Training Data in Federated Learning via Correlation Value Encoding and Segmented Aggregation
FedCVESA is a white-box attack on federated learning that demonstrates how a malicious server can encode and steal private training data from clients.
HAJJv2-CrowdCount: Zero-Shot Benchmark for Dense Crowd Counting
The HAJJv2-CrowdCount benchmark provides per-second human annotations for dense crowd counting and evaluates zero-shot counting paradigms like YOLO-World and SAM3Count.
Meta reuses old RAM in new servers with custom bridge chip
Meta is utilizing custom bridge chips to allow new servers to reuse older generations of RAM, extending hardware life and reducing costs.
AT-Attn: Temporal-Aware Cross-Attention for Longitudinal Multimodal Alzheimer's Disease Diagnosis
Researchers propose AT-Attn, a temporal-aware cross-attention framework for improving longitudinal multimodal diagnosis of Alzheimer's disease.
GeoProp: Grounding Robot State in Vision for Generalist Manipulation
GeoProp is introduced as a lightweight adapter that aligns robotic proprioception with vision through geometric grounding to improve generalist manipulation policies.
Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning
A Tree-of-Thoughts (ToT) reasoning framework is proposed for text-to-image in-context learning to reduce prompt ambiguity and compositional errors.
Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning
The Entropy Pacing Policy Optimization (EPPO) framework is introduced to stabilize multi-task agentic reinforcement learning for LLMs by coordinating entropy across tasks.
Predicting LLM Safety Before Release by Simulating Deployment
Researchers demonstrate a method for predicting LLM safety before release by simulating deployment using de-identified conversations from previous releases.
Validate the Dream Before You Trust Its Verdict: Admissibility for World-Model Simulators
A proposed 'admissibility ladder' framework aims to certify world-model simulators in robotics to ensure their verdicts are trustworthy for policy evaluation.
Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 -- A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvency
A quantitative scenario analysis explores the AI industry's future economics (2026-2030), focusing on memory scarcity, inference efficiency, and infrastructure solvency.
Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation
A scoping review analyzes the current state of vision foundation models in radiology, highlighting promising transferability but gaps in clinical translation.
DiPhon: Diffusion on Graphons for Scalable Graph Generation
DiPhon is introduced as a diffusion framework on graphons for size-scalable graph generation, allowing models to generate larger graphs without retraining.
Gimitest: A Comprehensive Tool for Testing Reinforcement Learning Policies
Gimitest is an open-source framework designed to test the reliability and safety of single- and multi-agent reinforcement learning policies across various gym environments.
AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning
AnchorPrune is a training-free framework for visual token pruning in vision-language models, improving inference efficiency by preserving query-critical evidence and complementary context.