← All posts / Models

8B Open Weights Beat the Closed Frontier: SenseTime's SenseNova-U1.5 Outscores Nano-Banana-Pro on Vision Reasoning

SenseTime's 8B-parameter SenseNova-U1.5-8B-MoT natively unifies visual understanding and generation without encoders or VAEs — and posts 68.2% on VBVR-Pro-Bench, ahead of proprietary Nano-Banana-Pro at 56.4%, under Apache 2.0.

8B Open Weights Beat the Closed Frontier: SenseTime's SenseNova-U1.5 Outscores Nano-Banana-Pro on Vision Reasoning

For most of the last two years, the deal in image AI was simple: closed models made the prettier pictures, and open models made the cheaper ones. That deal just got torn up. SenseTime’s SenseNova-U1.5 — an 8B-parameter natively unified multimodal model released under Apache 2.0 — not only tops the open-source leaderboard on generation and editing benchmarks; on VBVR-Pro-Bench, a reasoning-intensive vision evaluation, it scores 68.2% against 56.4% for Google’s proprietary Nano-Banana-Pro. An open checkpoint you can run yourself just beat a frontier closed model at its own game.

What SenseNova-U1.5 Actually Is

The “U” series is SenseTime’s SenseNova take on the unified-multimodal problem, and its distinguishing move is architectural radicalism. Most multimodal systems are Frankenstein builds: a vision encoder (often a CLIP or SigLIP variant) digests images into embeddings, a separate diffusion decoder with a VAE generates pixels, and glue code shuttles between them. SenseNova-U1.5 dispenses with all of it. The model is encoder-free and VAE-free — a Mixture-of-Transformers (MoT) that operates directly on visual patches in the same token space it uses for language.

The payoff is that understanding and generation share one representation. When the model renders text into a poster, the same weights that read posters are at work. The architecture builds on the NEO-unify design described in SenseTime’s earlier SenseNova-U1 paper (arXiv:2605.12500); U1.5 strengthens the patchify layers, tightens data quality and distribution, improves task formulation, and upgrades the prompt-enhancement and post-training pipeline. Two checkpoints ship on Hugging Face: the RL-trained main model and an intermediate SFT version for researchers who want to inspect the training ladder.

The Numbers

The headline results from the official release:

  • VBVR-Pro-Bench (reasoning-intensive vision): 68.2% — above Nano-Banana-Pro at 56.4%
  • GenEval (compositional generation): 0.92 — best among open-source models
  • CVTG-2K (text rendering in images): 0.948
  • ImgEdit (instruction-based editing): 4.59
  • GEdit-Bench-EN: top open-source score
  • Native 4K generation out of a single 8B-class model

Two of these deserve emphasis. Text rendering inside generated images has long been the embarrassment of diffusion models — the reason every AI-generated signpost looked like it was written by someone having a stroke. CVTG-2K at 0.948 across Chinese and English means legible posters, infographics, and brand assets are now a solved-enough problem at the open-weights tier. And the VBVR-Pro-Bench margin matters most of all: it’s not an aesthetics score, it measures whether a model can reason about what it sees. Beating a closed frontier product by nearly twelve points there reframes what “open-source” means in this category.

Six Improvements in the Official Release

SenseTime is unusually specific about what changed from the preview checkpoint:

  1. Higher-quality generation — better composition, color harmony, material rendering, and finer local detail
  2. Better text rendering and infographics — clearer information hierarchy in text-dense designs
  3. More efficient native 4K generation — coherent global structure at high resolution
  4. More reliable image editing — stronger identity and background preservation across local edits, insertions, and multi-reference replacements
  5. Stronger complex-instruction following — consistent handling of object counts, spatial relations, and multiple constraints per request
  6. More precise visual control — region- and object-level steering via bounding boxes, visual markers, and reference images

Just as notable is the honesty of the “Ongoing Improvements” section, which concedes real weaknesses: oversaturated colors on some prompts (mitigable by lowering cfg_scale), errors in dense mixed Chinese-English text, imperfect exact-count layouts, unstable small faces and hands, and drift on broad multi-region edits. That candor is rare in model releases and useful for anyone planning production deployment.

Why It Matters

The open frontier is closing the gap from below. Qwen, DeepSeek, and now SenseTime have spent 2026 demonstrating that open weights can sit at or near the frontier on specific axes rather than merely trailing it. SenseNova-U1.5 is the cleanest example yet in visual creation: an 8B model — small enough to fine-tune on a single node — outreasoning a proprietary Google product. For startups, that removes the platform risk of building on someone else’s rate-limited API. For researchers, an Apache 2.0 RL checkpoint with its SFT sibling is a gift: you can study the entire post-training pipeline, not just the endpoint.

Efficiency is the quiet story. No VAE and no separate encoder means fewer components, less memory, and a shorter path from tokens to pixels. The model card reports “improved generation efficiency” for native 4K — and 4K output from an 8B-class unified model would have sounded absurd eighteen months ago, when 4K meant cascaded upscalers bolted onto multi-billion-parameter diffusion stacks.

The unified paradigm is winning. SenseNova-U1.5, alongside contemporaries like BAGEL and Qwen’s unified efforts, marks the field’s drift away from pipeline architectures toward single models that see and draw with the same brain. The VBVR-Pro-Bench result suggests this isn’t just elegant engineering — shared representations may genuinely help visual reasoning.

Try It

Weights are on Hugging Face (sensenova/SenseNova-U1.5-8B-MoT) and ModelScope, with a reference implementation in the SenseNova-U1 GitHub repo (Python 3.11, PyTorch 2.8, CUDA 12.8). Inference for text-to-image and editing ships as example scripts; a free browser playground, SenseNova-Studio, requires no GPU. BF16 weights total roughly 35GB sharded, and quantizations are already available.

The broader takeaway: if your product roadmap assumed open image models would stay a generation behind, it’s time to re-run that assumption. The gap didn’t just narrow — on reasoning-heavy vision benchmarks, it inverted.