All Articles
16650 articles total
Gritt exits stealth with $34 million for robots to build solar plants—then, everything else
Robotics startup Gritt emerges from stealth with $34 million in funding to automate construction site tasks, specifically for solar plants.
Who’s afraid of the big, bad GPU?
An exploration of the environmental and social costs of the AI GPU boom, including energy consumption, water usage, and e-waste.
WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement
Introduction of WuYu-EnvLE-Bench, a benchmark for evaluating LLMs in environmental law enforcement cases.
Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models
MIND, a cognitive jailbreak framework for text-to-image models that models latent defense mechanisms to generate adversarial prompts.
Financial Audit Assistance using Misinformation Detection and Explanation
A system using unsupervised techniques for misinformation detection and explanation in financial audit assistance.
PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model
PGN (Pangu Navigator), an offline vision-language navigation system based on the OpenPangu-7B multimodal foundation model.
A Hardware-oriented Approach for Efficient Bayesian Inference Computation and Deployment
A hardware-oriented approach to accelerate discrete Bayesian inference on embedded GPUs via memory layout restructuring and tensor clustering.
Exploratory and Assimilating Reflection: Reflective Recall Cycle for Long-term Memory
The EAR framework introduces Exploratory and Assimilating Reflection to improve long-term memory retrieval for LLM-based autonomous agents.
ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding
ST-Veto, a training-free method to improve diffusion multimodal LLM reasoning by vetoing temporally unstable and weakly grounded tokens.
Mechanistic Attention Guidance for Agent Memory Refinement
Researchers propose AGMR, a framework that uses retrieval-head attention signals to guide targeted memory updates for AI agents, improving efficiency and performance over text-only methods.
Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents
The VRR-Stop framework introduces a robust stopping mechanism for verify-repair loops in LLM agents using a noise model and belief filtering to prevent damaging correct plans.
FlowBlock: Wavefront-Parallel Decoding for Self-Correcting Diffusion Language Models
FlowBlock introduces wavefront-parallel decoding for self-correcting diffusion language models, significantly increasing tokens per second and reducing latency without requiring retraining.
OrientSAM: Mitigating Camera-Centric Shortcut in Multimodal Spatial Reasoning via Orientation-Aware Spatial Alignment
OrientSAM is a framework designed to improve multimodal spatial reasoning in LLMs by injecting explicit orientation information via Fourier-based angle encoding.
Artificial Intelligence for Understanding and Managing Transportation Behavior in Sustainable Smart Cities
A study explores the application of AI for managing urban transportation behavior in smart cities, emphasizing a behavior-centered perspective on mobility data.
ProEvent: An Event-centric Benchmark for Proactive Agents
ProEvent is a new event-centric benchmark to evaluate the ability of proactive AI agents to maintain user timetables from chat interactions, revealing current LLM limitations.
LaT: LLM-as-Trainer for Multi-Task Vehicle Routing Solvers
The LLM-as-Trainer (LaT) paradigm uses a pretrained LLM as an external trainer to provide stage-wise guidance for multi-task vehicle routing solvers.
Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution
Researchers analyze cross-modal negation detection in vision-language models and propose an attention-based architecture to better model inter-modal dependencies.
SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategies Refinement in E-Commerce Recommendation
SR-Agent is an experience-driven framework for automating the refinement of post-ranking strategies in industrial e-commerce recommender systems.
Semantically Similar, Logically Distinct: Diagnosing the Semantic-Answerability Gap in Table RAG
The TCR-Bench benchmark diagnoses the 'Semantic-Answerability Gap' in Table RAG, showing that semantic relevance does not guarantee that a table can actually answer a query.
Qwen-Image-3.0: Rich Content, Authentic Details, Deep Knowledge
Alibaba introduces Qwen-Image-3.0, a multimodal model focusing on high-fidelity image understanding and detailed knowledge extraction.