All Articles
17562 articles total
Server-side Anti-cheat in FPS games for Aimbot detection using Deep learning and Machine learning
Development of YAACS, a server-side aimbot detection system for FPS games using Stacked LSTM and deep learning to minimize false positives.
Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs
Introduction of Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super that significantly increases server throughput and concurrency using a hybrid MoE pruning approach.
Decentralized Aggregation of LLM Predictions via Wagering Mechanisms
Proposal of WALLA, a decentralized wagering mechanism for aggregating LLM predictions that ensures incentive compatibility and advantage-aligned weights.
MechMath Agent Team: LLM Driven Agents for Mathematical Research
The MechMath Agent Team (MMAT) is a multi-agent LLM system designed to co-pilot mathematical research and produce formally certified proofs.
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL
LLM-as-a-Tutor is a framework for non-verifiable RL that adapts prompt difficulty in real-time based on the evolving capabilities of the policy.
Agent Step Value: State-Transition Measurement with State-Grounded LLM Evaluators
Introduction of Agent Step Value (ASV), a framework for measuring the impact of individual agent actions on state transitions using LLM evaluators.
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
ResearchStudio-Idea is a skill suite for ML research ideation that uses evidence grounding and a corpus of conference papers to generate traceable research proposals.
PLACEMEM: Toward a Compute-Aware Memory Plane for Lifelong Agents
The authors propose PLACEMEM, a compute-aware memory plane for lifelong agents that uses versioned capsules to manage persistence, evolution, and correction of agent memories.
Forethought: Verifiable Reasoning from Neurosymbolic Primitive Programming
Forethought is a neurosymbolic reasoning system that treats reasoning as explicit, verifiable programs, allowing small models to match the capabilities of frontier models.
Language models guide symbolic equation discovery by controlling search
The paper introduces LLM-PySR, a framework where language models control the search process of symbolic regression to discover scientific equations more effectively.
A Clustering-Based Framework for Identifying Suspicious Trading Patterns in Capital Market
A study implements an unsupervised fraud-detection toolkit using K-Means++ clustering to identify suspicious trading patterns in capital markets.
Agentic IoT: Architectures, Applications, and Challenges Toward the Internet of Agents
This paper outlines the concept of Agentic IoT, moving from passive data collection to distributed cognitive agent ecosystems across the device-to-cloud continuum.
Unsupervised Features Mining via Activation Geometry
The Mining via Activation Geometry (MAG) framework extracts unsupervised reasoning features from LLM activations to enable vector steering and improved data selection.
Biological Motifs for Agentic Control
This research maps biological control motifs to agentic software design patterns using a typed interface correspondence to improve reliability and security in LLM agents.
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning
ProGPO is introduced as a learned-critic-free method for context-consistent step-level reinforcement learning in agentic tasks.
Shortcut Learning in Legal Judgment Prediction: Empirical Evidence from the UK Employment Tribunal
An empirical study reveals that legal judgment prediction models often rely on 'shortcut learning' from outcome-revealing cues in post-hoc judicial texts.
Agentic SABRE: An Uncertainty-Aware Neuro-Symbolic Multi-Agent Framework for Adaptive Ransomware Detection
Agentic SABRE is a neuro-symbolic multi-agent framework for adaptive ransomware detection that combines semantic evidence with behavioral telemetry and uncertainty quantification.
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry
PivoARL is a new self-feedback retry framework for LLM agents that identifies pivotal errors to reduce redundant interactions and improve learning efficiency.
Explainable Reinforcement Learning for Adaptive Traffic Signal Control
Researchers propose an explainable RL framework for traffic signal control that uses entity-centric architecture and attention networks for transparent decision-making.
Can Conversational Temporal Dynamics Improve Depression Detection in Dyads? A Preliminary Investigation in Multi-Modality Perspectives
A study demonstrates that conversational temporal dynamics (turn-pair timing) can serve as a lightweight and interpretable modality for improving depression detection in dyads.