All Articles
19202 articles total
An Integrated System for Real-Time Student Assessment and Career Guidance Using Neural Networks in Computing Disciplines
An AI-driven career guidance system for CS students utilizing a Multilayer Perceptron model and a web-based assessment platform.
AIChilles: Automatically Uncovering Hidden Weaknesses in AI-Evolved Systems
Introduces AIChilles, a tool designed to automatically uncover performance and correctness regressions in AI-evolved system programs.
Heteroskedastic Signals in Budgeted LLM Verification: Structural Heterogeneity Limits Optimization Gains
Analyzes how structural heterogeneity in uncertainty signals limits the gains of global optimization in budgeted LLM verification.
RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
Introduces RetailBench, a simulation benchmark for evaluating the long-horizon reasoning and decision-making capabilities of LLM agents in retail environments.
STRIDE: Strategic Trajectory Reasoning via Discriminative Estimation for Verifiable Reinforcement Learning
Proposes STRIDE, a fine-grained RLVR framework that uses verifiable outcomes to improve credit assignment and reasoning in LLMs.
Malaysia’s AI agent-powered messaging app Respond.io raises $62.5M, eyes acquisitions
Respond.io, a Malaysian AI-powered messaging app for customer inquiries, has raised $62.5M for expansion and potential acquisitions.
Advanced Machine Learning and Deep Learning Techniques for Enhanced Cattle Identification and Detection: A Comprehensive Review
A systematic review exploring the use of machine learning and deep learning techniques, such as CNNs and YOLO, for cattle identification and detection in livestock management.
Overcoming the Impedance Mismatch: A Theoretical Roadmap for Fusing Foundation Models and Knowledge Graphs
A theoretical paper addressing the 'Impedance Mismatch' between foundation models and knowledge graphs, proposing a roadmap for true semantic fusion using structured residual streams.
Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
Introduction of WebStep, a benchmark for process-level evaluation of web agents using semantic state tracking to pinpoint specific skill failures.
Multi-agent Framework for Time-Sensitive Complementary Collaboration in Minecraft
Presentation of TickingCollabBench and the TickingCollab framework for evaluating multi-agent complementary collaboration in time-sensitive Minecraft environments.
Recurrent Reasoning on Symbolic Puzzles with Sequence Models
Introduction of RecurrReason, a difficulty-controlled benchmark for testing recurrent logic puzzles on sequence models to analyze reasoning robustness.
Do LLMs Reliably Identify Correct Information Units in Aphasic Discourse?
A study evaluating the ability of LLMs like Llama-3.1 and Qwen2.5 to identify Correct Information Units (CIUs) in aphasic discourse transcripts.
Artificial Intelligence Index Report 2026
The 2026 AI Index Report provides a comprehensive overview of AI's progress in reasoning, safety, economy, and its impact on science and medicine.
AI-Driven Framework for Adaptive Water Network Management with Proof-of-Concept Implementation: Addressing Non-Revenue Water in Jordan
A framework for adaptive water network management in Jordan combining EPANET, digital twins, and local LLMs (via Ollama) to reduce non-revenue water loss.
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
Introduction of RoboPIN and Pinned Chain-of-Thought, a paradigm that anchors reasoning steps to visual evidence to improve embodied AI grounding.
John Carmack on Fabrice Bellard
A discussion on Hacker News featuring perspectives from John Carmack regarding the work of Fabrice Bellard.
Chili peppers of the world: cultivars, species, and heat
A community discussion about the various cultivars and heat levels of chili peppers worldwide.
QoS-Aware Token Scheduling and Private Data Valuation for Multi-Modal Agentic Networks
Research on fair token allocation and private data valuation in decentralized multi-modal agentic networks using differentially private prototypes.
Do we have the knowledge we need? Rethinking human-AI decision-making in corporations
A position paper exploring how organizations should maintain knowledge for both human and AI accessibility and how to allocate agency between them.
Large Language Models as Optimizers: A Survey of Direct vs. Tool-Augmented Approaches and Their Performance Frontiers
A survey of LLMs as optimizers, analyzing direct, tool-augmented, and tool-creating optimization paradigms.