Modern multimodal models integrate generation and understanding into a unified system, enabling learning from self-generated feedback. Introducing Uni Evo-VL, a self-evolving framework that leverages self-critiques as privileged information during test-time compute. Instead of relying on a separate, often larger teacher model, a single multimodal model acts as both teacher and student with different contexts.
The student only sees the vanilla question, while the teacher conditions on the privileged critique. Training minimizes the per-state divergence between their denoising diffusion distributions over the student's sampling trajectories. Experiments demonstrate that Uni Evo-VL improves image generation capabilities while maintaining sensitivity to reflection information.
Building on the open-source Qwen-image-2512, performance gains were observed from 0.747 to 0.808 on Gen Eval and from 32.97 to 35.53 on Gen Eval2 Soft-TIFA. Attempts with more powerful external critics, such as GPT5.6-Luna, show that multimodal models with strong judge capabilities can anticipate a higher self-evolving ceiling. Mixed text-rendering outcomes indicate that self-improvements may not be uniform across different tasks.
The study aims to enhance user experience when using multimodal models without external supervision or guidance. It sheds light on the current hot recursive self-improvement research line.
Source: Hacker News · Summarized by HeadlinesBriefing