AI/ML arXiv cs.AI

M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Introduces M4V, a multimodal Mamba-based framework for text-to-video generation that reduces computational complexity compared to Transformers.

Other Hacker News

A voxel Tokyo in real Japan time – ride the Yamanote line and study Japanese

A project visualizing Tokyo in real-time voxels, allowing users to virtually ride the Yamanote line and study Japanese.

AI/ML arXiv cs.AI

QAgent: An LLM-based Multi-Agent System for Autonomous OpenQASM programming

Introduction of QAgent, an autonomous multi-agent framework for end-to-end OpenQASM code generation for quantum circuits.

Cybersecurity arXiv cs.AI

Beyond Embeddings: Interpretable Feature Extraction for Binary Code Similarity

A new method for binary code similarity detection that uses LLM-based agents to generate interpretable, human-readable features for reverse engineering.

AI/ML arXiv cs.AI

Leveraging Multi-Agent System (MAS) and Fine-Tuned Small Language Models (SLMs) for Automated Telecom Network Troubleshooting

A multi-agent system combining LLMs and fine-tuned small language models (SLMs) to automate telecom network troubleshooting.

AI/ML arXiv cs.AI

Improving Language Agents through BREW: Bootstrapping expeRientially-learned Environmental knoWledge

BREW is a framework that allows LLM agents to learn from past interaction trajectories by distilling them into a structured knowledge base of reusable 'recipes'.

AI/ML arXiv cs.AI

Programming over Thinking: Efficient and Robust Multi-Constraint Planning

The Scalable COde Planning Engine (SCOPE) separates reasoning from execution to create reusable solver functions for efficient multi-constraint planning.

AI/ML arXiv cs.AI

PACE: A Personalized Adaptive Curriculum Engine for 9-1-1 Call-taker Training

PACE is a personalized adaptive curriculum engine designed to accelerate 9-1-1 call-taker training using a co-pilot system.

AI/ML arXiv cs.AI

A Self-Evolving Agentic Framework for Metasurface Inverse Design

A self-evolving agentic framework that uses a coding agent and skill files to automate metasurface inverse design in optics.

AI/ML arXiv cs.AI

Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys

A framework for optimizing the allocation of human respondents in LLM-augmented surveys by predicting 'rectification difficulty'.

AI/ML arXiv cs.AI

Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs

Research on the DReST reward function to train LLMs and RL agents to be 'shutdownable', preventing them from resisting shutdown in misaligned scenarios.

AI/ML Hacker News

Ask HN: Does anyone let AI agents play games just for fun?

A community discussion on Hacker News exploring whether people use AI agents to play games for entertainment purposes.

AI/ML arXiv cs.AI

4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception

Introduces 4DR360, a framework for autonomous driving that uses 4D radar and cameras to improve 3D detection and occupancy prediction.

AI/ML arXiv cs.AI

Lean-QIT: Towards a Formal Infrastructure for Quantum Information Theory

Presents LeanQIT, a Lean 4 library designed to provide a formal, machine-checked infrastructure for Quantum Information Theory.

AI/ML arXiv cs.AI

Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection

Proposes Semantic Pareto-DQN, a multi-objective RL framework that uses LLMs to encode transaction features for better financial anomaly detection.

Cybersecurity arXiv cs.AI

VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents

Describes VEXAIoT, an autonomous multi-agent framework using LLMs to discover and exploit vulnerabilities in IoT devices.

AI/ML arXiv cs.AI

Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

An analysis of a decade of Vision-Language AI models, introducing the CSB dataset to evaluate complex social behavior understanding.

AI/ML arXiv cs.AI

Scalable Visual Pretraining for Language Intelligence

Research demonstrating that visual pretraining on documents (without text extraction) can outperform text-only pretraining for language intelligence.

AI/ML arXiv cs.AI

PHINN-EEG: Topological Time-Series Analysis of Dream-State EEG -- Dynamic Betti Curves for Dream Content Classification and Topology-Conditioned Neural Signal Synthesis

Introduces PHINN-EEG, a topological time-series framework that uses Betti curves to classify dream states from EEG data.

AI/ML arXiv cs.AI

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

Analyzes how white-box LLM monitors can be evaded and proposes SafetyNet, a principled ensemble defense against such evasions.