All Articles
18727 articles total
The Null Is Always False (Except When It Is True) (2014)
A discussion or post regarding the nuances of null values and their truthiness in programming.
Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments
Research on the mismatch between AI tutor benchmarks and real-world usage, finding that students often bypass pedagogical scaffolding to reach their goals.
Mitigating Visual Hallucinations in Multimodal Systems through Retrieval-Augmented Reliability-Aware Inference
Proposes a retrieval-augmented reliability-aware inference framework to reduce visual hallucinations in multimodal LLMs without requiring retraining.
Unassigned Agents in Compilation-based Multi-agent Path Finding
Introduces UA-MAPF, a variant of multi-agent path finding with unassigned agents, and demonstrates its implementation using SMT-CBS and NRF-SAT solvers.
TrustedARI: Towards Trust-Native Agentic Routing Infrastructure for Agentic AI
Introduces TrustedARI, a trust-native routing infrastructure for AI agents featuring a custom TLS handshake, privacy-preserving query construction, and verifiable billing.
An Integrated System for Real-Time Student Assessment and Career Guidance Using Neural Networks in Computing Disciplines
An AI-driven career guidance system for CS students utilizing a Multilayer Perceptron model and a web-based assessment platform.
AIChilles: Automatically Uncovering Hidden Weaknesses in AI-Evolved Systems
Introduces AIChilles, a tool designed to automatically uncover performance and correctness regressions in AI-evolved system programs.
Heteroskedastic Signals in Budgeted LLM Verification: Structural Heterogeneity Limits Optimization Gains
Analyzes how structural heterogeneity in uncertainty signals limits the gains of global optimization in budgeted LLM verification.
RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
Introduces RetailBench, a simulation benchmark for evaluating the long-horizon reasoning and decision-making capabilities of LLM agents in retail environments.
STRIDE: Strategic Trajectory Reasoning via Discriminative Estimation for Verifiable Reinforcement Learning
Proposes STRIDE, a fine-grained RLVR framework that uses verifiable outcomes to improve credit assignment and reasoning in LLMs.
Malaysia’s AI agent-powered messaging app Respond.io raises $62.5M, eyes acquisitions
Respond.io, a Malaysian AI-powered messaging app for customer inquiries, has raised $62.5M for expansion and potential acquisitions.
Advanced Machine Learning and Deep Learning Techniques for Enhanced Cattle Identification and Detection: A Comprehensive Review
A systematic review exploring the use of machine learning and deep learning techniques, such as CNNs and YOLO, for cattle identification and detection in livestock management.
Overcoming the Impedance Mismatch: A Theoretical Roadmap for Fusing Foundation Models and Knowledge Graphs
A theoretical paper addressing the 'Impedance Mismatch' between foundation models and knowledge graphs, proposing a roadmap for true semantic fusion using structured residual streams.
Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
Introduction of WebStep, a benchmark for process-level evaluation of web agents using semantic state tracking to pinpoint specific skill failures.
Multi-agent Framework for Time-Sensitive Complementary Collaboration in Minecraft
Presentation of TickingCollabBench and the TickingCollab framework for evaluating multi-agent complementary collaboration in time-sensitive Minecraft environments.
Recurrent Reasoning on Symbolic Puzzles with Sequence Models
Introduction of RecurrReason, a difficulty-controlled benchmark for testing recurrent logic puzzles on sequence models to analyze reasoning robustness.
Do LLMs Reliably Identify Correct Information Units in Aphasic Discourse?
A study evaluating the ability of LLMs like Llama-3.1 and Qwen2.5 to identify Correct Information Units (CIUs) in aphasic discourse transcripts.
Artificial Intelligence Index Report 2026
The 2026 AI Index Report provides a comprehensive overview of AI's progress in reasoning, safety, economy, and its impact on science and medicine.
AI-Driven Framework for Adaptive Water Network Management with Proof-of-Concept Implementation: Addressing Non-Revenue Water in Jordan
A framework for adaptive water network management in Jordan combining EPANET, digital twins, and local LLMs (via Ollama) to reduce non-revenue water loss.
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
Introduction of RoboPIN and Pinned Chain-of-Thought, a paradigm that anchors reasoning steps to visual evidence to improve embodied AI grounding.