All Articles
17178 articles total
M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Introduces M4V, a multimodal Mamba-based framework for text-to-video generation that reduces computational complexity compared to Transformers.
A voxel Tokyo in real Japan time – ride the Yamanote line and study Japanese
A project visualizing Tokyo in real-time voxels, allowing users to virtually ride the Yamanote line and study Japanese.
QAgent: An LLM-based Multi-Agent System for Autonomous OpenQASM programming
Introduction of QAgent, an autonomous multi-agent framework for end-to-end OpenQASM code generation for quantum circuits.
Beyond Embeddings: Interpretable Feature Extraction for Binary Code Similarity
A new method for binary code similarity detection that uses LLM-based agents to generate interpretable, human-readable features for reverse engineering.
Leveraging Multi-Agent System (MAS) and Fine-Tuned Small Language Models (SLMs) for Automated Telecom Network Troubleshooting
A multi-agent system combining LLMs and fine-tuned small language models (SLMs) to automate telecom network troubleshooting.
Improving Language Agents through BREW: Bootstrapping expeRientially-learned Environmental knoWledge
BREW is a framework that allows LLM agents to learn from past interaction trajectories by distilling them into a structured knowledge base of reusable 'recipes'.
Programming over Thinking: Efficient and Robust Multi-Constraint Planning
The Scalable COde Planning Engine (SCOPE) separates reasoning from execution to create reusable solver functions for efficient multi-constraint planning.
PACE: A Personalized Adaptive Curriculum Engine for 9-1-1 Call-taker Training
PACE is a personalized adaptive curriculum engine designed to accelerate 9-1-1 call-taker training using a co-pilot system.
A Self-Evolving Agentic Framework for Metasurface Inverse Design
A self-evolving agentic framework that uses a coding agent and skill files to automate metasurface inverse design in optics.
Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys
A framework for optimizing the allocation of human respondents in LLM-augmented surveys by predicting 'rectification difficulty'.
Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs
Research on the DReST reward function to train LLMs and RL agents to be 'shutdownable', preventing them from resisting shutdown in misaligned scenarios.
Ask HN: Does anyone let AI agents play games just for fun?
A community discussion on Hacker News exploring whether people use AI agents to play games for entertainment purposes.
4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception
Introduces 4DR360, a framework for autonomous driving that uses 4D radar and cameras to improve 3D detection and occupancy prediction.
Lean-QIT: Towards a Formal Infrastructure for Quantum Information Theory
Presents LeanQIT, a Lean 4 library designed to provide a formal, machine-checked infrastructure for Quantum Information Theory.
Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection
Proposes Semantic Pareto-DQN, a multi-objective RL framework that uses LLMs to encode transaction features for better financial anomaly detection.
VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents
Describes VEXAIoT, an autonomous multi-agent framework using LLMs to discover and exploit vulnerabilities in IoT devices.
Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models
An analysis of a decade of Vision-Language AI models, introducing the CSB dataset to evaluate complex social behavior understanding.
Scalable Visual Pretraining for Language Intelligence
Research demonstrating that visual pretraining on documents (without text extraction) can outperform text-only pretraining for language intelligence.
PHINN-EEG: Topological Time-Series Analysis of Dream-State EEG -- Dynamic Betti Curves for Dream Content Classification and Topology-Conditioned Neural Signal Synthesis
Introduces PHINN-EEG, a topological time-series framework that uses Betti curves to classify dream states from EEG data.
Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors
Analyzes how white-box LLM monitors can be evaded and proposes SafetyNet, a principled ensemble defense against such evasions.