
COMMUNITY PAGE
Run Ideogram V4 on Floyo
Home / Model / Ideogram V4 on Floyo
AI IMAGE GENERATION
Run Ideogram V4 on Floyo
The #1 open-weight image model for design. 9.3B parameter DiT with 0.97 OCR text accuracy, JSON-controlled layout with bounding boxes, hex color palette conditioning, and native 2K resolution. Trained from scratch on structured captions. LoRA training supported.
Run Ideogram's V4 through ComfyUI in your browser. No API key, no installs, no local GPU.
|
Parameters 9.3B (single-stream DiT) |
Text Accuracy 0.97 OCR (best open-weight) |
|
Resolution Native 2K (up to 2048x2048) |
Arena Rank #1 open-weight (Design Arena) |
No installation. Runs in browser. Updated July 2026.
ideogram v4
image editing
image to image
Upload an image and describe the change you want. Ideogram V4 transforms it while preserving your composition, with the strength slider controlling how far the result moves from the original.
Ideogram V4 · Image to Image
Upload an image and describe the change you want. Ideogram V4 transforms it while preserving your composition, with the strength slider controlling how far the result moves from the original.
custom style
fine-tune
ideogram v4
lora
text to image
Paste the URL of a trained Ideogram V4 LoRA and write a prompt. The workflow generates images with your custom subject, style, or identity applied, using up to three LoRAs at once.
Ideogram V4 + LoRA · Text to Image
Paste the URL of a trained Ideogram V4 LoRA and write a prompt. The workflow generates images with your custom subject, style, or identity applied, using up to three LoRAs at once.
custom style
fine-tune
ideogram v4
lora training
Upload your training images and this workflow fine-tunes a LoRA adapter on Ideogram V4, returning a download link for the trained LoRA file and config, ready to use in generation workflows.
Ideogram V4 LoRA Trainer · LoRA Training
Upload your training images and this workflow fine-tunes a LoRA adapter on Ideogram V4, returning a download link for the trained LoRA file and config, ready to use in generation workflows.
ideogram v4
text rendering
text to image
typography
Write a prompt and Ideogram V4 generates a design-grade image with the most accurate text rendering of any open-weight model, at up to native 2K resolution.
Ideogram V4 · Text to Image
Write a prompt and Ideogram V4 generates a design-grade image with the most accurate text rendering of any open-weight model, at up to native 2K resolution.
What you get?
Ideogram V4 is the #1 open-weight image model for design work, released June 3, 2026 by Ideogram AI. A 9.3 billion parameter single-stream Diffusion Transformer trained from scratch on structured JSON captions. Highest text rendering accuracy among open-weight models (0.97 X-Omni OCR) ahead of models 2-4x its size. Bounding-box layout control, hex color palette conditioning, native 2K resolution, and Qwen3-VL-8B vision-language text encoder. LoRA training and image-to-image supported. Available as ComfyUI nodes on Floyo with 4 workflows.
IDEOGRAM V4 WORKFLOWS ON FLOYO
_1782815066892_1783181186386.webp?width=1400&height=620&quality=80&resize=cover)
_1782815066892_1783181186386.webp?width=1400&height=620&quality=80&resize=cover)
_1782818117647_1783181192198.webp?width=1400&height=620&quality=80&resize=cover)
_1782818117647_1783181192198.webp?width=1400&height=620&quality=80&resize=cover)
_1782833857439_1783181250627.webp?width=1400&height=620&quality=80&resize=cover)
_1782833857439_1783181250627.webp?width=1400&height=620&quality=80&resize=cover)
_1782818163695_1783181256618.webp?width=1400&height=620&quality=80&resize=cover)
_1782818163695_1783181256618.webp?width=1400&height=620&quality=80&resize=cover)
_1782815066892_1783181186386.webp?width=104&height=104&quality=80&resize=cover)
_1782818117647_1783181192198.webp?width=104&height=104&quality=80&resize=cover)
_1782833857439_1783181250627.webp?width=104&height=104&quality=80&resize=cover)
_1782818163695_1783181256618.webp?width=104&height=104&quality=80&resize=cover)
What is Ideogram V4?
Ideogram V4 is Ideogram AI's first open-weight text-to-image model and its most capable release. Released June 3, 2026, it is a 9.3 billion parameter foundation model trained from scratch. Not a fine-tune. Not a distillation. Every parameter was trained from the ground up on structured JSON captions that describe composition, style, lighting, color, typography, and spatial layout per element.
The result is a model that feels more like a design tool than an image generator. Specify bounding-box coordinates to place a headline in the upper third. Set hex colors to lock a brand palette. Describe each text element with its literal string and a separate visual styling instruction. The model follows these as deterministic constraints, not suggestions. This level of structural control is what separates Ideogram V4 from everything else in the open-weight space.
Text rendering is Ideogram's signature capability, and V4 pushes it further than any open-weight model. It scores 0.97 on X-Omni English OCR accuracy, ahead of FLUX.2 dev (32B), HunyuanImage 3.0 (80B MoE), and Qwen-Image (20B). Multi-line headlines, varied font weights, logos, signage, captions, watermarks, and bilingual text all render legibly. In a blind typography evaluation by ContraLabs using ten professional designers, Ideogram V4 was rated highest among all open-weight models.
The text encoder choice is what makes this possible. Instead of CLIP or T5, Ideogram V4 uses Qwen3-VL-8B-Instruct, a full vision-language model. Hidden states from 13 of its intermediate layers are concatenated, giving the model multi-scale semantic features from surface-level token understanding to deep compositional reasoning. This is why it handles dense, multi-clause design briefs better than models using text-only encoders.
On Floyo, Ideogram V4 runs through native ComfyUI nodes on H100 NVL GPUs. Four workflows cover text-to-image, image-to-image, LoRA-powered generation, and LoRA training. No model downloads, no 24GB GPU requirement, no JSON prompt formatting needed (the workflows handle it).
What are Ideogram V4's technical specifications?
Ideogram V4 is a 9.3B parameter single-stream Diffusion Transformer with 34 layers, 4608 embedding dimensions, 18 attention heads, and 128 latent channels. Text encoder is Qwen3-VL-8B-Instruct (vision-language model with 13-layer hidden state extraction). Trained on structured JSON captions with bounding-box layout and hex color conditioning. Native 2K resolution. Scores 0.97 OCR, 0.89 prompt alignment, 0.76 spatial reasoning.
| Spec | Details |
|---|---|
| Developer | Ideogram AI (Toronto) |
| Architecture | Single-stream flow-matching DiT (34 layers, 4608 embed dim, 18 heads) |
| Parameters | 9.3 billion |
| Text Encoder | Qwen3-VL-8B-Instruct (vision-language model, 13-layer hidden state concatenation) |
| Training Data | Structured JSON captions (per-element styling, bounding boxes, color palettes) |
| Max Resolution | 2048x2048 (native 2K) |
| Latent Channels | 128 |
| Max Text Tokens | 2,048 |
| Attention | QK-RMSNorm + 3D MRoPE |
| MLP | SwiGLU blocks with AdaLN scale/gate timestep conditioning |
| X-Omni OCR (EN) | 0.97 (best open-weight) |
| Prompt Alignment (Prism) | 0.89 |
| Layout Control (7Bench) | 0.69 (better than all closed-source models) |
| Design Arena | #1 open-weight model (Elo ~1285) |
| Internal Benchmark | #2 overall (behind GPT Image 2 only) |
| Quantizations | FP8, NF4 (runs on 24GB consumer GPU) |
| LoRA Support | Yes (training + inference) |
| Weights License | Non-commercial (self-hosting). Commercial via API or enterprise license. |
| ComfyUI Access | Native support on Floyo (4 workflows) |
| Release Date | June 3, 2026 |
What can you create with Ideogram V4?
Ideogram V4 covers poster design, ad creatives, product packaging, brand identity assets, social media graphics, UI mockups, signage, menus, infographics, editorial imagery, fashion lookbooks, bilingual marketing content, and any work where legible text inside the image is a requirement. The JSON prompt system and bounding-box control make it closer to a design tool than a general image generator.
| Capability | What It Does | Use Case |
|---|---|---|
| Design-Grade Typography | 0.97 OCR accuracy. Multi-line headlines, varied font weights, logos, signage, captions, and watermarks render legibly. Specific typeface names outperform generic style descriptions. | Posters, ad creatives, packaging, branded assets |
| Bounding-Box Layout | Specify x/y coordinates to place subjects, text elements, and background regions. Overlapping boxes create z-order layering. Headlines stay where you put them. | Magazine layouts, multi-region compositions, precise placements |
| Color Palette Conditioning | Pass hex color codes to lock the image's dominant palette. Keeps output aligned to brand guidelines without post-processing color correction. | Brand-accurate campaigns, seasonal palettes, style guides |
| Image-to-Image | Upload a source image and apply transformations: restyle, relight, recompose, or add elements while preserving the core composition. | Product variations, style transfer, iterative refinement |
| LoRA Training + Inference | Train custom LoRAs on your own brand data. Fine-tune for specific product lines, character IP, or house styles. Generate with trained LoRAs loaded. | Brand consistency, product-specific models, enterprise deployment |
| Pipeline Integration | Chain with video models in ComfyUI. Generate with Ideogram V4, animate with Wan 2.7 or Vidu Q3, add voiceover with ElevenLabs. Or convert to 3D with TRELLIS 2. | Multi-model production pipelines |
How does Ideogram V4 compare to other image models?
Ideogram V4 leads all open-weight models on text rendering (0.97 OCR) and layout control (7Bench). GPT Image 2 leads on aesthetic Elo and overall quality. Nano Banana Pro leads on 4K native resolution and character consistency. FLUX.2 dev leads on raw parameter scale (32B). ERNIE Image leads on prompt enhancement. Ideogram V4's edge: design-grade typography, JSON-controlled layout, and the strongest text-per-parameter ratio in the field.
| Model | Parameters | Text (OCR) | Layout Control | Design Arena |
|---|---|---|---|---|
| Ideogram V4 | 9.3B | 0.97 | Bounding-box + JSON | #1 open-weight |
| GPT Image 2 | GPT-5.4 backbone | ~99% | Reasoning-based | #1 overall (~1405 Elo) |
| Nano Banana Pro | Gemini backbone | 94%+ | Thinking mode | Top tier |
| FLUX.2 dev | 32B | Moderate | Via Kontext | Top tier |
| ERNIE Image | 8B | 0.9733 LTB | Prompt enhancer | N/A |
Source: Design Arena leaderboard, X-Omni OCR benchmark, 7Bench layout control, ContraLabs blind designer evaluation, Ideogram internal Bradley-Terry benchmark, and HuggingFace model cards as of July 2026.
What are Ideogram V4's key features?
Ideogram V4's feature set is engineered for design work, not general image generation. The JSON caption training, bounding-box control, color palette conditioning, and vision-language text encoder all target the same goal: make the model behave like a design tool that follows a brief, not a creative AI that interprets a mood.
#1 Open-Weight Text Rendering (0.97 OCR)
Ideogram has led on in-image typography since its first release. V4 pushes further: 0.97 X-Omni OCR accuracy, the highest among all open-weight models. This is ahead of FLUX.2 dev (32B), HunyuanImage 3.0 (80B MoE), and Qwen-Image (20B). Dense multi-line text, varied font weights, logos with taglines, bilingual signage, and eight-item tour date posters all render cleanly. This is the feature that makes Ideogram V4 irreplaceable for design work.
Structured JSON Prompting
Every training image had exhaustively described JSON captions: per-element styling, literal text strings with separate visual descriptions, bounding-box coordinates, and color palettes. The model understands text placement at a structural level. Double-quoted text renders cleaner than unquoted. Specific typeface names (Cooper Black, Futura, Playfair Display) outperform generic descriptions. Adding "professional typography" to the style field pulls from a higher-quality training slice.
Bounding-Box Spatial Control
Specify x/y bounding-box coordinates in your prompt to place subjects, text, logos, and background regions precisely. Overlapping boxes create z-order: later elements render on top of earlier ones. This is how to place a logo over a photograph without a separate mask layer. The model treats coordinates as deterministic constraints, not suggestions.
Color Palette Conditioning
Pass hex color codes in a color_palette array to steer the image's dominant color scheme. Brand guidelines translate directly into model parameters. Hex codes work in the array field; color words like "deep navy blue" outperform hex codes in prose descriptions. Both methods keep the output within your brand palette without post-production color grading.
Vision-Language Text Encoder (Qwen3-VL-8B)
Instead of CLIP or T5, Ideogram V4 uses Qwen3-VL-8B-Instruct as its text encoder. Hidden states from 13 intermediate layers are concatenated, giving multi-scale semantic features from surface token understanding to deep compositional reasoning. This is why the model handles dense, multi-clause design briefs better than larger models using traditional text-only encoders.
LoRA Training Pipeline
Train custom LoRAs directly on Floyo. Fine-tune V4 for your brand's product line, character IP, or house style. The trained LoRA loads into the generation workflow for consistent, on-brand output. Enterprises can fine-tune on their own data and deploy within their own infrastructure.
Transparency and Layered Output (Roadmap)
Background Remover is live today: clean alpha cutouts from any V4 output. Coming next: editable text layers and movable image layers directly from inference. The model will return output as a stack of components rather than a flat frame. Headlines can be revised after generation. This closes the gap between AI output and production-ready design files.
How does Ideogram V4 work?
Ideogram V4 is a flow-matching text-to-image model with a fully single-stream DiT architecture. Text and image tokens are concatenated into one unified sequence and processed through the same 34-layer transformer. The text encoder is Qwen3-VL-8B-Instruct, a vision-language model whose hidden states from 13 layers are concatenated to produce multi-scale semantic features for the DiT.
The model was trained exclusively on structured JSON captions. Every training image was described with per-element breakdowns: the literal text string, its visual styling, its spatial position (bounding box), the overall color palette (hex codes), and the compositional structure. This training approach is why the model treats layout instructions as deterministic constraints rather than vague style guidance.
For inference, the model supports multiple sampler presets. V4_QUALITY_48 is the highest-quality preset (48 steps). Default and Turbo presets trade quality for speed. NF4 quantization fits the full model on a 24GB consumer GPU. FP8 quantization offers higher quality at about 29.5GB. On Floyo's H100 NVL GPUs, the full model runs without quantization for maximum output quality.
On Floyo, Ideogram V4 runs through native ComfyUI nodes. The text-to-image workflow accepts your prompt (plain text or JSON structured), the image-to-image workflow adds a source image for transformation, the LoRA workflow loads your fine-tuned adapters, and the LoRA trainer handles the full training pipeline. All workflows run on H100 NVL GPUs.
Fair warning: Ideogram V4 open weights are under a non-commercial license for self-hosting. Commercial use requires the Ideogram API or enterprise license. On Floyo, this distinction does not apply because Floyo handles the model access. The model is strongest at design-grade work (posters, typography, packaging, branding). It trails GPT Image 2 on aesthetic Elo and photorealistic human portraiture. Hands can break in busy compositions. Multi-person scenes can get muddy.
Frequently Asked Questions
Common questions about running Ideogram V4 on Floyo.
You can start with Floyo's free pricing plan. To continue using the service beyond the free tier, upgrade your Floyo pricing plan. Ideogram V4 runs as a native ComfyUI node on Floyo's H100 NVL GPUs, so there is no additional API cost beyond your Floyo plan.
Open Floyo in your browser, search "Ideogram V4" in the template library, and pick a workflow (text-to-image, image-to-image, LoRA generation, or LoRA training). Click Run, write your prompt, and generate. Floyo handles the GPU, ComfyUI environment, and model weights. No 24GB GPU, no Python setup, no JSON formatting needed.
Ideogram AI, a research lab based in Toronto. Founded by Mohammad Norouzi (co-founder, CEO) and team members from Google Brain who created Imagen and co-authored the DDPM paper. Ideogram V4 was released June 3, 2026 as their first open-weight model. Weights are on HuggingFace. Code is on GitHub (ideogram-oss/ideogram4).
Two architectural choices. First: the text encoder is Qwen3-VL-8B-Instruct, a full vision-language model that understands text at a compositional level, not a keyword level. Second: the model was trained on structured JSON captions where every text element was described with its literal string, visual styling, and spatial position. The model learned text placement as a structural task, not a pattern-matching task.
Ideogram V4 leads on layout control (bounding boxes, JSON structure) and is the strongest open-weight model for design work. GPT Image 2 leads on overall aesthetic quality (Elo ~1405 vs ~1285), photorealistic portraiture, and batch consistency (8 images per prompt). For posters, packaging, branded assets, and typography-heavy work, Ideogram V4 is the pick. For general-purpose photorealism, GPT Image 2 has the edge. Both are available on Floyo.
Yes. The "Ideogram V4 LoRA Trainer" workflow runs the full training pipeline on Floyo's H100 NVL GPUs. Provide your training images, configure the adapter, and train. The output LoRA loads into the "Ideogram V4 LoRA Text to Image" workflow for generation. Fine-tune for brand products, character IP, or house styles.
Yes. Floyo runs ComfyUI, which lets you chain multiple models. Generate a design with Ideogram V4, animate with Wan 2.7 or Vidu Q3, add voiceover with ElevenLabs or Fish Audio S2, upscale with Topaz. Or generate product shots, then convert to 3D with TRELLIS 2 or Meshy v6. All in one pipeline.
On Floyo, yes. Floyo handles the model access. The open weights on HuggingFace carry a non-commercial license for self-hosting. Commercial use of the weights requires the Ideogram API ($0.03-$0.10/image), fal.ai, or a direct enterprise license. On Floyo, this licensing distinction does not affect you because Floyo manages the model access.
Try Ideogram V4 on Floyo
#1 open-weight image model for design. 9.3B parameters, 0.97 OCR text accuracy, bounding-box layout, color palette control, LoRA training, and native 2K. Run it in your browser.
Try Ideogram V4 Now → Browse All ModelsRelated Reading
AI Ad Creatives for Social and Web
Character and Concept Design on Floyo
Last updated: July 2026. Specs from Ideogram AI official press release (June 3, 2026), GitHub (ideogram-oss/ideogram4), HuggingFace model cards, Design Arena leaderboard, X-Omni OCR benchmark, 7Bench layout control, ContraLabs blind designer evaluation, and third-party reviews.
_1782818117647.png?width=400&height=300&quality=80&resize=cover)
_1782833857439.png?width=400&height=300&quality=80&resize=cover)

_1782815066892.png?width=400&height=300&quality=80&resize=cover)