All Articles
19072 articles total
MMLongEmbed: Benchmarking Multimodal Embedding Models in Long-Context Scenarios
MMLongEmbed is a new benchmark designed to evaluate Multimodal Embedding Models in long-context scenarios across text, document, and video.
Is My Vision-Language Data in Your AI? Membership Inference Test (MINT) Demo 2
The Membership Inference Test (MINT) Demo 2 provides a framework and web platform to detect if specific data were used to train AI models.
Automated 3D Kinematic Monitoring for Circadian Activity and Anomaly Detection in Juvenile Fish
A new 3D behavioral phenotyping framework combines deep learning and stereo vision to monitor circadian activity and stress in juvenile tilapia.
Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech
Pixel-TTS is a framework that renders text as images to improve robustness and zero-shot generalization in text-to-speech synthesis.
X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
X-Tokenizer introduces a multimodal action tokenizer using Semantic Residual Quantization to provide a shared interface for robot control in VLA models.
Beyond Self-Attention: Sub-Quadratic Vision Transformers for Fast Image Captioning
A new sub-quadratic Vision Transformer replaces self-attention with a GMM-based probabilistic approach to speed up image captioning.
Sub-Semantic Image Segmentation
The DETECTURE framework enables sub-semantic image segmentation by coupling VLMs with SAM 3 to partition images into stable appearance patterns.
Where Does Texture Evidence Live in SAM? Features, Proposal Masks, and Texture Segmentation
Research explores whether frozen Segment Anything Models (SAM) possess internal texture evidence and how they fail at texture-defined partitions.
ASM SHADER TOY – It's shader toy but you code in asm
ASM Shader Toy is a creative coding platform that allows users to write shaders using assembly language.
Frood, an Alpine Initramfs NAS
Frood is an Alpine Linux-based Initramfs NAS designed for lightweight storage management.
Show HN: VoiceDraw – Talk system design out loud, the diagrams draw themselves
VoiceDraw is a tool that automatically generates system design diagrams from spoken descriptions in real-time.
MiroBench: Benchmarking Realism in Agentic Simulation of Real-world Discussions
The MiroBench benchmark evaluates the realism of LLM agents simulating real-world Reddit discussions across semantic and structural dimensions.
RAMS: Resource-Adaptive and Detection-Conditioned Model Switching for Embedded Edge Perception
RAMS is a runtime controller for embedded edge perception that dynamically switches between YOLOv8 model tiers to optimize latency and accuracy.
Gender Differences in AI Literacy Workshop Outcomes and Deepfake Engagement
A study on Australian students explores how gender affects AI literacy, deepfake engagement, and STEM career aspirations.
VigilFormer: Deformable Attention for Video Anomaly Detection with Causal Risk Inference
VigilFormer is a video anomaly detection framework using deformable spatio-temporal attention and causal risk inference for high-speed surveillance.
Steady-Forcing: Balancing Spatial Persistence and Motion Continuity in Long-Horizon Nature Video Diffusion
Steady-Forcing improves long-horizon nature video diffusion by balancing spatial persistence and motion continuity using a visual anchor and motion memory.
BRIDGE: Biological Evidence Refinement and Heterogeneous Dynamic Gating for Gene Regulatory Networks
BRIDGE is a framework for gene regulatory network inference from scRNA-seq data using contrastive learning and heterogeneous dynamic gating.
Do Large Language Models Have Emotions?
This paper critiques the claim that LLMs have 'functional emotions' by comparing internal representations to biological affective neuroscience.
Nobody clicks your share buttons
A discussion on the inefficiency and low click-through rates of social media share buttons on websites.
Databricks Launches LTAP: A Unified OLAP/OLTP Data Architecture
Databricks introduces LTAP, a unified data architecture aiming to merge OLAP and OLTP capabilities.