All Articles
18697 articles total
Disentangling Hallucinations: Orthogonal Semantic Projection for Robust Interpretability
The OSP method aims to reduce semantic hallucinations in Vision-Language Models by orthogonalizing query vectors against distractor concepts.
Temporally Consistent and Controllable Video Generation of 2D Cine CMR via Latent Space Motion Modeling
A new generative framework for synthesizing temporally coherent 2D Cine CMR cardiac sequences using latent space motion modeling.
GeoRoPE: Ground-Aware Rotary Adaptation for Remote Sensing Foundation Models
GeoRoPE is a ground-aware spatial adaptation method for remote sensing foundation models to improve cross-resolution robustness.
GrapheneOS has been ported to Android 17 and official releases are coming soon
GrapheneOS has been ported to Android 17, with official releases expected soon.
Z.ai’s open-weights GLM-5.2 beats GPT-5.5 on multiple long-horizon coding benchmarks for 1/6th the cost
Z.ai released GLM-5.2, an open-weights 753B parameter model with a 1M token context window and 'IndexShare' optimization for autonomous coding.
MMLongEmbed: Benchmarking Multimodal Embedding Models in Long-Context Scenarios
MMLongEmbed is a new benchmark designed to evaluate Multimodal Embedding Models in long-context scenarios across text, document, and video.
Is My Vision-Language Data in Your AI? Membership Inference Test (MINT) Demo 2
The Membership Inference Test (MINT) Demo 2 provides a framework and web platform to detect if specific data were used to train AI models.
Automated 3D Kinematic Monitoring for Circadian Activity and Anomaly Detection in Juvenile Fish
A new 3D behavioral phenotyping framework combines deep learning and stereo vision to monitor circadian activity and stress in juvenile tilapia.
Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech
Pixel-TTS is a framework that renders text as images to improve robustness and zero-shot generalization in text-to-speech synthesis.
X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
X-Tokenizer introduces a multimodal action tokenizer using Semantic Residual Quantization to provide a shared interface for robot control in VLA models.
Beyond Self-Attention: Sub-Quadratic Vision Transformers for Fast Image Captioning
A new sub-quadratic Vision Transformer replaces self-attention with a GMM-based probabilistic approach to speed up image captioning.
Sub-Semantic Image Segmentation
The DETECTURE framework enables sub-semantic image segmentation by coupling VLMs with SAM 3 to partition images into stable appearance patterns.
Where Does Texture Evidence Live in SAM? Features, Proposal Masks, and Texture Segmentation
Research explores whether frozen Segment Anything Models (SAM) possess internal texture evidence and how they fail at texture-defined partitions.
ASM SHADER TOY – It's shader toy but you code in asm
ASM Shader Toy is a creative coding platform that allows users to write shaders using assembly language.
Frood, an Alpine Initramfs NAS
Frood is an Alpine Linux-based Initramfs NAS designed for lightweight storage management.
Show HN: VoiceDraw – Talk system design out loud, the diagrams draw themselves
VoiceDraw is a tool that automatically generates system design diagrams from spoken descriptions in real-time.
MiroBench: Benchmarking Realism in Agentic Simulation of Real-world Discussions
The MiroBench benchmark evaluates the realism of LLM agents simulating real-world Reddit discussions across semantic and structural dimensions.
RAMS: Resource-Adaptive and Detection-Conditioned Model Switching for Embedded Edge Perception
RAMS is a runtime controller for embedded edge perception that dynamically switches between YOLOv8 model tiers to optimize latency and accuracy.
Gender Differences in AI Literacy Workshop Outcomes and Deepfake Engagement
A study on Australian students explores how gender affects AI literacy, deepfake engagement, and STEM career aspirations.
VigilFormer: Deformable Attention for Video Anomaly Detection with Causal Risk Inference
VigilFormer is a video anomaly detection framework using deformable spatio-temporal attention and causal risk inference for high-speed surveillance.