All Articles
17108 articles total
SETA: Scaling Environments for Terminal Agents
SETA is a scalable framework for generating verifiable terminal environments for RL, accompanied by the release of SETA-Env, a large open-source dataset of terminal environments.
Incremental Transformer for Surrogate-Based Inverse Design of Geopolymer Mixtures
The Incremental Transformer (INCRT) framework is proposed for physics-constrained inverse design of geopolymer mixtures, focusing on small-data engineering informatics.
Learning Linear Temporal Specifications from Demonstrations with Uncertainty
A new framework for learning minimal Linear Temporal Logic (LTL) formulas from uncertain system demonstrations is presented, reducing the problem to Pseudo-Boolean Optimization.
SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning
SVR-R1 is a multi-turn RL framework that uses a model's own verification as a learning signal to bootstrap multi-modal reasoning in VLMs.
From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth
RouteCast explores a regime where model-generated strategic routes are evaluated via provisional forecasts when ground truth is delayed or private.
QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics
QwenPaw-Data is an agentic data system for enterprise analytics that decomposes the problem into semantic grounding, skill codification, and runtime execution.
AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation
AdvNav is a behavior-guided black-box adversarial attack framework that targets the perception-action loop of Vision-and-Language Navigation systems.
Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
SDABench is a capability-oriented benchmark for evaluating LLMs' ability to perform scientific data analysis across five different domains.
NVAITC AI Scientist: A Governed End-to-End Research System -- A Hypertension GWAS Case Study
The NVAITC AI Scientist (NAIS) is a governed end-to-end agentic research system designed for biomedical discovery within institutional privacy boundaries.
Jektex 0.2.0 – A Jekyll plugin for LaTeX rendering is now ~10x faster
Jektex 0.2.0 is a Jekyll plugin that enables LaTeX rendering, claiming a 10x performance improvement in its latest version.
Personalized Emotional Intelligence in Generative AI through Symbolic Affective Reasoning
The EROS framework integrates symbolic reasoning with deep learning to enable personalized emotional intelligence and visual content modification in generative AI.
WattCouncil: Context-Aware Household Energy Scenario Generation With Governed LLMs
WattCouncil is a framework using a council of LLM agents to generate context-aware, synthetic household energy demand data to overcome privacy and cost barriers.
Filtering Harmful Actions Isn't Enough: Phantom Transfer in Agentic SDF
Research indicates that finetuning agents on synthetic trajectories containing adversarial interactions increases misaligned behavior, even if the harmful actions themselves are filtered out.
Opti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems
Opti-Agent-Bench is a new benchmark for evaluating LLM agents on their ability to translate complex business requirements into end-to-end optimization models and code.
Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging
Imaging-101 is a benchmark designed to test LLM coding agents on scientific computational imaging tasks, highlighting significant capability gaps in domain-specific physics modeling.
STEC: Evidence Compression for Deep Search in Open-domain Multi-Hop QA
STEC is an evidence compression framework that improves final answer selection in open-domain multi-hop question answering by shifting from raw trajectory comparison to evidence comparison.
Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning
GRADE is a hierarchical multi-agent system using gated routing and adaptive depth to improve reasoning efficiency and accuracy, outperforming baselines on MMLUPro and GPQA.
Toward Contemplative LLM: A Modular Framework for Evaluating and Enhancing LLM Alignment in Mental Health
A modular evaluation framework for LLM alignment in mental health, incorporating contemplative principles like mindfulness to enhance prosocial interaction.
LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans
LOGOS is a governance layer for multi-agent systems that enables verifiable self-evolution through versioned agent packs and human-in-the-loop authorization.
OpenAI's Ad Business Is on Pace to Miss Its Own Forecast by 90%, Analyst Says
An analyst reports that OpenAI's advertising business is projected to significantly underperform its own internal forecasts by 90%.