OmniVChat audio-visual dialogue paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“OmniVChat offers a fresh approach for audio-visual dialogues with AI, cutting out text queries and speech recognition. This improves interaction speed and uses synthetic dialogues in a robust evaluation framework that aligns AI responses with real-world sce…”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“A new spectral deflation approach improves matrix filtering in the Muon optimizer and semidefinite programming. This method preserves key spectral components and enhances accuracy, resulting in better GPT-2 pretraining and lower errors in large-scale optimi…”
ArenaFlow agent RL paper
2 directory members surfaced this signal.
2 experts
1 community
1 sources clustered
“Qiang Zhang, Ruixue Ding, Fanrui Zhang, Xi Chen, Boli Chen, Shihang Wang, Yinfeng Huang, Yi Zheng, Pengjun Xie, Kaipeng Zhang, Jiawei Liu, Zheng-Jun Zha ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL https://ar…”
“Alibaba's ArenaFlow enhances open-ended reinforcement learning via credit propagation, optimizing performance through pivotal reasoning and skill cultivation. This addresses reward discrimination collapse, leading to improved sample efficiency in complex sc…”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Survival Reinforcement Learning (SRL) outperforms existing methods by 2x to 8x on long-horizon tasks, showcasing the potential of online classification strategies for scaling self-supervised RL, promising greater stability in complex dynamical systems. http…”
IntTravel travel recommendation dataset
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Amap unveils IntTravel, a groundbreaking dataset and framework transforming travel recommendations by integrating key journey aspects—departure time, travel mode, and on-the-way needs. This innovation, linked to a 1.09% rise in CTR, reshapes mobility for mi…”
RAVE visual attention multimodal paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“RAVE improves visual attention in large multimodal models, addressing issues of misallocation between text and images. This mechanism yields an average 3-point boost on perception-heavy tasks, enhancing precise and reliable visual grounding in AI. https://a…”
RecreationWorld computer-use agents paper
2 directory members surfaced this signal.
2 experts
1 community
1 sources clustered
“Shuai Bai, Jiayong Deng, Yikun Fu, Chang Gao, Xuhao Hu, Mianqiu Huang, Yizhen Jiang, Yuheng Jing, Dehui Kong, Keliang Li, Ning Li, Wanli Li, Dayiheng Liu, Dunjie Lu, ... RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents ht…”
“RECREATIONWORLD merges GUI exploration with code creation in hybrid computer-use agents, enabling them to autonomously recreate applications across five platforms. This fosters self-improvement through verifiable feedback, pushing automation limits. https:/…”
WorldRoamBench open-world models benchmark
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“WorldRoamBench evaluates interac- tive world models with a benchmark for long-horizon stability in action, visual quality, physics, and memory. No model excels in all areas, revealing the need for robust solutions in virtual simulations. https://arxiv.org/a…”
PrismAlign VLM table OCR paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“PrismAlign enhances table OCR by aligning various Vision-Language Models to minimize hallucinations and boost extraction accuracy. Utilizing a novel Bayesian approach, it achieves top performance in benchmarks, tackling persistent document understanding cha…”
AI-GRACE agentic deployment framework paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“AI-GRACE helps organizations operationalize agentic AI by linking governance with technical implementation, ensuring trust, risk management, and value realization. This framework streamlines deployment while enhancing safety and accountability. https://arxi…”
Omni Demand multimodal intent benchmark paper
2 directory members surfaced this signal.
2 experts
1 community
1 sources clustered
“Qi Chen, Yunfei Chu, Haolin He, Yifan Yang, Zihan Liu, Yuxuan Wang, Ziyang Ma, Ruiyang Xu, Meng Gao, Yinsong Yan, Ling Wang, Hui Wang, Wen Huang, ... Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction https…”
“This study outlines Omni Demand Understanding (ODU), evaluating AI's capacity to interpret user intent in complex multimodal interactions. Current models lack context understanding, exposing a key gap in intelligent assistance that could reshape AI communic…”
diffusion language models parallelism paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“New research reveals that uniform and Gaussian diffusion enable efficient sampling with significantly fewer forward passes than masked diffusion, reshaping our understanding of model parallelism and optimizing generative AI techniques. https://arxiv.org/abs…”
L0 MoE dense LLM acceleration paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“L0-MoE speeds up large language models by 2.5x with minimal performance loss, leveraging L0-regularization and a cluster confusion matrix for efficient training, enhancing LLM accessibility against current costly methods. https://arxiv.org/abs/2609.21672”
IntBMoE block-level MoE paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“DreamX launched IntBMoE, a Mixture-of-Experts architecture that decouples expert participation, execution, and materialization for optimal expert use and lower compute costs. Its deployment in Alibaba’s AMap system has driven strong performance gains for mi…”
dense MoE vision language action paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“AdaDE enhances vision-language-action policies by converting dense networks into efficient MoE structures, enabling dynamic expert deactivation that retains 95.7% success in complex robotic tasks while reducing model size by 40%. https://arxiv.org/abs/2609.…”
2 directory members surfaced this signal.
2 experts
1 community
1 sources clustered
“Haodi Lei, Yafy Li, Haoran Zhang, Shunkai Zhang, Qianjia Cheng, Xiaoye Qu, Ganqu Cui, Bowen Zhou, Ning Ding, Yun Luo, Yu Cheng Draft-OPD: On-Policy Distillation for Speculative Draft Models https://arxiv.org/abs/2605.29343”
“This study introduces Draft-OPD, an on-policy distillation method that boosts speculative decoding efficiency in large language models. Leveraging target-assisted rollouts, Draft-OPD achieves over 5× speedup without loss, outperforming EAGLE-3 and DFlash. h…”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“This study introduces Bayesian Chronicle Agents (BCA), enhancing LLM simulations with a transparent belief layer for controlled opinion dynamics. This approach recovers stubbornness in behavior and reveals biases, making simulations more realistic and audit…”
CompAdapt text-to-video motion paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Introducing CompAdapt, a physics-consistent framework for text-to-video generation that boosts realism through complex motion dynamics, including collisions and coupled movements. With one-shot adaptation to new laws, it exceeds traditional models in visual…”