[Try Seedance 2.5 in fal Agent](/content/agent?endpoint=bytedance%2Fseedance-2.5%2Fimage-to-video/index.html)

[Docs](/content/docs/documentation/index.html)

[Log-in](/content/login?returnTo=/explore/index.html) [Sign-up](/content/login?returnTo=/explore/index.html)

[Contact Sales](/content/enterprise#contact-sales/index.html) [Log-in](/content/login?returnTo=/explore/index.html) [Sign-up](/content/login?returnTo=/explore/index.html)

new

image-to-video

stylized

transform

lipsync

# Seedance 2.5 Image to Video

Dreamina Seedance 2.5 animates a single still into a native 30-second clip at up to 720p, extending one frame into continuous, coherent motion without the drift or stitching of shorter multi-clip workflows.

[Try it now!](/content/models/bytedance/seedance-2.5/image-to-video/index.html) [See docs](/content/models/bytedance/seedance-2.5/image-to-video/api/index.html)

new

image-to-video

stylized

transform

lipsync

# MiniMax H3

MiniMax H3 is a frontier video model. This endpoint animates a supplied image into 2K video, using it as the opening frame or pairs a first and last frame to control a transition between two images with the aspect ratio following the input.

[Try it now!](/content/models/minimax/h3/image-to-video/index.html) [See docs](/content/models/minimax/h3/image-to-video/api/index.html)

new

image-to-video

stylized

transform

lipsync

# Flux 3

FLUX.3 is Black Forest Labs' frontier video model. This endpoint animates a single still image into video, extending one frame into coherent, natural motion.

[Try it now!](/content/models/blackforestlabs/flux-3/image-to-video/index.html) [See docs](/content/models/blackforestlabs/flux-3/image-to-video/api/index.html)

![Kling Video v3 Image to Video [Pro]](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a8cfd08%2FJi4e0i6Afbeql3Wr5UTz6_ab60b14661424612bf19059e97e996a5.jpg/tr:w-1920,q-80/Ji4e0i6Afbeql3Wr5UTz6_ab60b14661424612bf19059e97e996a5.webp)

image-to-video

image-to-video

# Kling Video v3 Image to Video \[Pro\]

Kling 3.0 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.

[Try it now!](/content/models/fal-ai/kling-video/v3/pro/image-to-video/index.html) [See docs](/content/models/fal-ai/kling-video/v3/pro/image-to-video/api/index.html)

text-to-image

stylized

transform

typography

# Krea 2 Turbo

Generate high-fidelity images from text in seconds with Krea 2 Turbo, the speed-optimized open-source version of Krea 2, preserving its aesthetic range for rapid ideation.

[Try it now!](/content/models/fal-ai/krea-2/turbo/index.html) [See docs](/content/models/fal-ai/krea-2/turbo/api/index.html)

Seedance 2.5 Image to Video

MiniMax H3

Flux 3

Kling Video v3 Image to Video \[Pro\]

Krea 2 Turbo

Search models...Search by model, task, category and more

View all models

Try: [Newest image to video models](/content/explore/search?q=newest%20image%20to%20video%20models/index.html) [Flux Kontext](/content/explore/search?q=flux%20kontext/index.html) [Generate 3D model](/content/explore/search?q=generate%203d%20model/index.html) [Create music](/content/explore/search?q=create%20music/index.html) [Remove background](/content/explore/search?q=remove%20background/index.html) [Upscale](/content/explore/search?q=upscale/index.html) [Training](/content/explore/search?q=training/index.html) [Try on clothing](/content/explore/search?q=try%20on%20clothing/index.html)

[**Trending**](/content/explore/trending/index.html)

Models that are popular with developers right now.

Per section28

\\
\\
\\
nano-banana-2/edit\\
\\
Nano Banana 2 is Google's new state-of-the-art image generation and editing model\\
\\
image-to-image](/content/models/fal-ai/nano-banana-2/edit/index.html)

\\
\\
\\
nano-banana-pro/edit\\
\\
Nano Banana Pro is Google's new state-of-the-art image generation and editing model\\
\\
realism\\
\\
typography\\
\\
image-to-image](/content/models/fal-ai/nano-banana-pro/edit/index.html)

\\
\\
\\
openai/gpt-image-2/edit\\
\\
GPT Image 2, OpenAI's latest image model, is capable of making fine-grained, detailed edits to images.\\
\\
gpt-image-2\\
\\
openai\\
\\
chatgpt-images-2\\
\\
image-to-image](/content/models/openai/gpt-image-2/edit/index.html)

[![FLUX.1 [schnell] is a 12 billion parameter flow transformer that generates high-quality images from text in 1 to 4 steps, suitable for personal and commercial use.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a9af64d%2FGxvCUPd3gO-MSYcy06g0x_1641cfe028c2429b8e12e4fc320eb0a8.jpg/tr:w-1920,q-80/GxvCUPd3gO-MSYcy06g0x_1641cfe028c2429b8e12e4fc320eb0a8.webp)\\
\\
\\
flux/schnell\\
\\
FLUX.1 \[schnell\] is a 12 billion parameter flow transformer that generates high-quality images from text in 1 to 4 steps, suitable for personal and commercial use.\\
\\
text-to-image](/content/models/fal-ai/flux/schnell/index.html)

\\
\\
\\
nano-banana-2\\
\\
Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model\\
\\
text-to-image](/content/models/fal-ai/nano-banana-2/index.html)

\\
\\
\\
openai/gpt-image-2\\
\\
GPT Image 2, OpenAI's latest image model, is capable of creating extremely detailed images with fine typography.\\
\\
gpt-image-2\\
\\
openai\\
\\
typography\\
\\
text-to-image](/content/models/openai/gpt-image-2/index.html)

\\
\\
\\
nano-banana-pro\\
\\
Nano Banana Pro is Google's new state-of-the-art image generation and editing model\\
\\
realism\\
\\
typography\\
\\
text-to-image](/content/models/fal-ai/nano-banana-pro/index.html)

[![FLUX.1 [dev] is a 12 billion parameter flow transformer that generates high-quality images from text. It is suitable for personal and commercial use.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffal_cdn%2Ffal%2FUpscale-1.jpeg/tr:w-1920,q-80/Upscale-1.webp)\\
\\
\\
flux/dev\\
\\
FLUX.1 \[dev\] is a 12 billion parameter flow transformer that generates high-quality images from text. It is suitable for personal and commercial use.\\
\\
text-to-image](/content/models/fal-ai/flux/dev/index.html)

\\
\\
\\
nano-banana/edit\\
\\
Google's famous original image generation and editing model\\
\\
image-editing\\
\\
image-to-image](/content/models/fal-ai/nano-banana/edit/index.html)

\\
\\
\\
kling-video/v3/pro/image-to-video\\
\\
Kling 3.0 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.\\
\\
image-to-video](/content/models/fal-ai/kling-video/v3/pro/image-to-video/index.html)

\\
\\
\\
bytedance/seedance-2.0/image-to-video\\
\\
ByteDance's most advanced image-to-video model. Animate still images into cinematic video with synchronized audio, start and end frame control, and motion prompts.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/bytedance/seedance-2.0/image-to-video/index.html)

[![Image editing with FLUX.2 [pro] from Black Forest Labs. Ideal for high-quality image manipulation, style transfer, and sequential editing workflows](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2Fpenguin%2FUfryXXm9my6IM8HsoP9FL_054c2c2953dc491996904114c6e04836.jpg/tr:w-1920,q-80/UfryXXm9my6IM8HsoP9FL_054c2c2953dc491996904114c6e04836.webp)\\
\\
\\
flux-2-pro\\
\\
Image editing with FLUX.2 \[pro\] from Black Forest Labs. Ideal for high-quality image manipulation, style transfer, and sequential editing workflows\\
\\
text-to-image](/content/models/fal-ai/flux-2-pro/index.html)

\\
\\
\\
bytedance/seedance-2.0/reference-to-video\\
\\
ByteDance's most advanced reference-to-video model. Generate video from up to 9 images, 3 videos, and 3 audio clips with native audio and cinematic camera control.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/bytedance/seedance-2.0/reference-to-video/index.html)

\\
\\
\\
bytedance/seedream/v5/pro/edit\\
\\
Seedream 5.0 Pro is grounded, region-precise image editing model that changes one element while keeping the rest of the frame intact with layer separation, sketch completion, and up to 10 reference images.\\
\\
realism\\
\\
typography\\
\\
stylized\\
\\
image-to-image](/content/models/bytedance/seedream/v5/pro/edit/index.html)

\\
\\
birefnet/v2\\
\\
bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)\\
\\
background removal\\
\\
segmentation\\
\\
high-res\\
\\
image-to-image](/content/models/fal-ai/birefnet/v2/index.html)

[![FLUX.1 Kontext [pro] handles both text and reference images as inputs, seamlessly enabling targeted, local edits and complex transformations of entire scenes.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffal_cdn%2Ffal%2FTraining-2.jpg/tr:w-1920,q-80/Training-2.webp)\\
\\
\\
flux-pro/kontext\\
\\
FLUX.1 Kontext \[pro\] handles both text and reference images as inputs, seamlessly enabling targeted, local edits and complex transformations of entire scenes.\\
\\
image-to-image](/content/models/fal-ai/flux-pro/kontext/index.html)

[![FLUX1.1 [pro] is an enhanced version of FLUX.1 [pro], improved image generation capabilities, delivering superior composition, detail, and artistic fidelity compared to its predecessor.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffalserverless%2Fgallery%2Fturbo_thumbnail.jpg/tr:w-1920,q-80/turbo_thumbnail.webp)\\
\\
\\
flux-pro/v1.1\\
\\
FLUX1.1 \[pro\] is an enhanced version of FLUX.1 \[pro\], improved image generation capabilities, delivering superior composition, detail, and artistic fidelity compared to its predecessor.\\
\\
text-to-image](/content/models/fal-ai/flux-pro/v1.1)

\\
\\
\\
nano-banana\\
\\
Google's famous original image generation and editing model\\
\\
image-generation\\
\\
text-to-image](/content/models/fal-ai/nano-banana/index.html)

\\
\\
\\
kling-video/v2.5-turbo/pro/image-to-video\\
\\
Kling 2.5 Turbo Pro: Top-tier image-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.\\
\\
stylized\\
\\
transform\\
\\
image-to-video](/content/models/fal-ai/kling-video/v2.5-turbo/pro/image-to-video/index.html)

\\
\\
\\
bytedance/seedream/v4.5/edit\\
\\
A new-generation image creation model ByteDance, Seedream 4.5 integrates image generation and image editing capabilities into a single, unified architecture.\\
\\
stylized\\
\\
transform\\
\\
image-to-image](/content/models/fal-ai/bytedance/seedream/v4.5/edit/index.html)

\\
\\
seedvr/upscale/image\\
\\
Use SeedVR2 to upscale your images\\
\\
upscale\\
\\
image-to-image](/content/models/fal-ai/seedvr/upscale/image/index.html)

[![Text-to-image generation with FLUX.2 [pro] from Black Forest Labs. Optimized for maximum quality, exceptional photorealism and artistic images.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2Fpenguin%2FeZetcrsZI6AQLCD3f5gaI_b173ae004bdd4108bd1be54eb6e49c7a.jpg/tr:w-1920,q-80/eZetcrsZI6AQLCD3f5gaI_b173ae004bdd4108bd1be54eb6e49c7a.webp)\\
\\
\\
flux-2-pro/edit\\
\\
Text-to-image generation with FLUX.2 \[pro\] from Black Forest Labs. Optimized for maximum quality, exceptional photorealism and artistic images.\\
\\
image-to-image](/content/models/fal-ai/flux-2-pro/edit/index.html)

\\
\\
\\
kling-video/v3/standard/image-to-video\\
\\
Kling 3.0 Standard: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.\\
\\
image-to-video](/content/models/fal-ai/kling-video/v3/standard/image-to-video/index.html)

\\
\\
\\
bria/background/remove\\
\\
Bria RMBG 2.0 enables seamless removal of backgrounds from images, ideal for professional editing tasks. Trained exclusively on licensed data for safe and risk-free commercial use. Model weights for commercial use are available here: https://share-eu1.hsforms.com/2GLpEVQqJTI2Lj7AMYwgfIwf4e04\\
\\
background removal\\
\\
image segmentation\\
\\
high resolution\\
\\
image-to-image](/content/models/fal-ai/bria/background/remove/index.html)

[![FLUX1.1 [pro] ultra is the newest version of FLUX1.1 [pro], maintaining professional-grade image quality while delivering up to 2K resolution with improved photo realism.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffalserverless%2Fgallery%2Fflux-pro-v1-1-ultra.webp/tr:w-1920,q-80/flux-pro-v1-1-ultra.webp)\\
\\
\\
flux-pro/v1.1-ultra\\
\\
FLUX1.1 \[pro\] ultra is the newest version of FLUX1.1 \[pro\], maintaining professional-grade image quality while delivering up to 2K resolution with improved photo realism.\\
\\
high-res\\
\\
realism\\
\\
text-to-image](/content/models/fal-ai/flux-pro/v1.1-ultra/index.html)

\\
\\
\\
bytedance/seedream/v5/pro/text-to-image\\
\\
ByteDance's Seedream 5.0 Pro is flagship text-to-image model, with deep-thinking prompt understanding, native text in 14 languages, and precise control over dense layouts and structured designs.\\
\\
realism\\
\\
typography\\
\\
stylized\\
\\
text-to-image](/content/models/bytedance/seedream/v5/pro/text-to-image/index.html)

\\
\\
\\
xai/grok-imagine-image/edit\\
\\
Edit images precisely with xAI's Grok Imagine model\\
\\
grok\\
\\
xai\\
\\
image-editing\\
\\
image-to-image](/content/models/xai/grok-imagine-image/edit/index.html)

[![Text-to-image generation with FLUX.2 [klein] 9B from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a8a7f3c%2F90FKDpwtSCZTqOu0jUI-V_64c1a6ec0f9343908d9efa61b7f2444b.jpg/tr:w-1920,q-80/90FKDpwtSCZTqOu0jUI-V_64c1a6ec0f9343908d9efa61b7f2444b.webp)\\
\\
\\
flux-2/klein/9b\\
\\
Text-to-image generation with FLUX.2 \[klein\] 9B from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.\\
\\
text-to-image](/content/models/fal-ai/flux-2/klein/9b/index.html)

[**Recently Added**](/content/explore/recently-added/index.html)

Newly added models across image, video, audio, and more.

\\
\\
new\\
\\
\\
bytedance/seedance-2.5/reference-to-video\\
\\
Dreamina Seedance 2.5 generates video from up to 50 multimodal references images, video, audio, and style inputs, locking a character, set, and palette across a full 30-second take for production-grade consistency.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/bytedance/seedance-2.5/reference-to-video/index.html)

\\
\\
new\\
\\
\\
bytedance/seedance-2.5/image-to-video\\
\\
Dreamina Seedance 2.5 animates a single still into a native 30-second clip at up to 720p, extending one frame into continuous, coherent motion without the drift or stitching of shorter multi-clip workflows.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/bytedance/seedance-2.5/image-to-video/index.html)

\\
\\
new\\
\\
\\
bytedance/seedance-2.5/text-to-video\\
\\
Dreamina Seedance 2.5 generates native 30-second single-shot video at up to 720p from a single text prompt, reasoning about the whole shot at once so motion, lighting, and subject identity stay coherent from first frame to last.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-video](/content/models/bytedance/seedance-2.5/text-to-video/index.html)

\\
\\
new\\
\\
hitem3d/hi3d/texture\\
\\
Texture an existing geometry mesh using a reference image with Hi3D.\\
\\
3d-to-3d](/content/models/hitem3d/hi3d/texture/index.html)

\\
\\
new\\
\\
hitem3d/hi3d/image-to-relief\\
\\
Generate a 3D relief depth map with Hi3D from a single image.\\
\\
depth\\
\\
hi3d\\
\\
relief\\
\\
image-to-image](/content/models/hitem3d/hi3d/image-to-relief/index.html)

\\
\\
new\\
\\
hitem3d/hi3d/multicolor\\
\\
Convert a textured 3D model into a multicolor model suited for multicolor 3D printing with Hi3D.\\
\\
multicolor\\
\\
3d-to-3d](/content/models/hitem3d/hi3d/multicolor/index.html)

\\
\\
new\\
\\
hitem3d/hi3d/split\\
\\
Split a 3D model into parts with Hi3D.\\
\\
3d-to-3d](/content/models/hitem3d/hi3d/split/index.html)

\\
\\
new\\
\\
hitem3d/hi3d/multi-view-to-3d\\
\\
Generate 3D models from multiple view images using Hi3D.\\
\\
multiview-to-3d\\
\\
3d\\
\\
image-to-3d](/content/models/hitem3d/hi3d/multi-view-to-3d/index.html)

\\
\\
new\\
\\
hitem3d/hi3d/image-to-3d\\
\\
Generate 3D models from a single image with Hi3D.\\
\\
3d\\
\\
mesh\\
\\
image-to-3d](/content/models/hitem3d/hi3d/image-to-3d/index.html)

\\
\\
new\\
\\
tripo3d/tripo/segment\\
\\
Automatically splits a 3D model into semantic parts for editing, texturing, and rigging.\\
\\
stylized\\
\\
transform\\
\\
3d-to-3d](/content/models/tripo3d/tripo/segment/index.html)

\\
\\
new\\
\\
tripo3d/tripo/remesh\\
\\
Converts triangle meshes into clean quad topology at a target polygon count, animation-ready with no manual retopology.\\
\\
stylized\\
\\
transform\\
\\
3d-to-3d](/content/models/tripo3d/tripo/remesh/index.html)

\\
\\
new\\
\\
\\
bytedance/seedream/v5/pro/layerize\\
\\
Splits a finished image into independent, editable transparent-PNG layers — background plus separate elements, from a text description, returning 2 to 17 layers per call for non-destructive reuse in design tools.\\
\\
utility\\
\\
editing\\
\\
image-to-image](/content/models/bytedance/seedream/v5/pro/layerize/index.html)

\\
\\
new\\
\\
\\
blackforestlabs/flux-3/text-to-video\\
\\
FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-video](/content/models/blackforestlabs/flux-3/text-to-video/index.html)

\\
\\
new\\
\\
\\
blackforestlabs/flux-3/image-to-video\\
\\
FLUX.3 is Black Forest Labs' frontier video model. This endpoint animates a single still image into video, extending one frame into coherent, natural motion.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/blackforestlabs/flux-3/image-to-video/index.html)

\\
\\
new\\
\\
\\
blackforestlabs/flux-3/first-last-frame-to-video\\
\\
FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates the video between a defined start and end frame, interpolating a smooth, coherent transition from the first image to the last.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/blackforestlabs/flux-3/first-last-frame-to-video/index.html)

\\
\\
new\\
\\
\\
blackforestlabs/flux-3/keyframes-to-video\\
\\
FLUX.3 is Black Forest Labs' frontier video model. This endpoint builds video from a sequence of keyframes, generating the motion between each anchor point for precise control over how a shot progresses.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/blackforestlabs/flux-3/keyframes-to-video/index.html)

\\
\\
new\\
\\
\\
blackforestlabs/flux-3/extend-video\\
\\
FLUX.3 is Black Forest Labs' frontier video model. This endpoint continues an existing clip beyond its final frame, generating additional footage that stays consistent with the original motion and scene.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
video-to-video](/content/models/blackforestlabs/flux-3/extend-video/index.html)

\\
\\
new\\
\\
\\
blackforestlabs/flux-3/text-to-video/draft\\
\\
FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews from a text prompt, with a reusable draft cache for full-quality enhancement.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-video](/content/models/blackforestlabs/flux-3/text-to-video/draft/index.html)

\\
\\
new\\
\\
\\
blackforestlabs/flux-3/image-to-video/draft\\
\\
FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews that animate a still image, with a reusable draft cache for full-quality enhancement.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/blackforestlabs/flux-3/image-to-video/draft/index.html)

\\
\\
new\\
\\
\\
blackforestlabs/flux-3/first-last-frame-to-video/draft\\
\\
FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews between a start and an end frame, with a reusable draft cache for full-quality enhancement.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/blackforestlabs/flux-3/first-last-frame-to-video/draft/index.html)

\\
\\
new\\
\\
\\
blackforestlabs/flux-3/keyframes-to-video/draft\\
\\
FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews pinned to your keyframe images, with a reusable draft cache for full-quality enhancement.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/blackforestlabs/flux-3/keyframes-to-video/draft/index.html)

\\
\\
new\\
\\
\\
blackforestlabs/flux-3/extend-video/draft\\
\\
FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews that continue an existing clip, with a reusable draft cache for full-quality enhancement.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
video-to-video](/content/models/blackforestlabs/flux-3/extend-video/draft/index.html)

\\
\\
new\\
\\
\\
blackforestlabs/flux-3/draft-enhance\\
\\
FLUX.3 is Black Forest Labs' frontier audio/video model. Re-render a previously generated draft at full quality — same seed, same motion, no re-planning.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
video-to-video](/content/models/blackforestlabs/flux-3/draft-enhance/index.html)

\\
\\
new\\
\\
heygen/v3/filler-word-removal\\
\\
Use Heygen's Latest Model for Filler Word Removal.\\
\\
filler-word-removal\\
\\
video-to-video](/content/models/fal-ai/heygen/v3/filler-word-removal/index.html)

\\
\\
new\\
\\
\\
xai/grok-imagine-video/v1.5/reference-to-video\\
\\
Generate videos from images and audio references using xAI's Grok Imagine 1.5 Video model.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/xai/grok-imagine-video/v1.5/reference-to-video/index.html)

\\
\\
new\\
\\
\\
xai/grok-imagine-video/v1.5/text-to-video\\
\\
Generate videos from prompts with audio using xAI's Grok Imagine 1.5 Video model.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-video](/content/models/xai/grok-imagine-video/v1.5/text-to-video/index.html)

\\
\\
new\\
\\
\\
minimax/h3/image-to-video\\
\\
MiniMax H3 is a frontier video model. This endpoint animates a supplied image into 2K video, using it as the opening frame or pairs a first and last frame to control a transition between two images with the aspect ratio following the input.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/minimax/h3/image-to-video/index.html)

\\
\\
new\\
\\
\\
minimax/h3/text-to-video\\
\\
MiniMax H3 is a frontier video model. This endpoint generates video from a text prompt alone, rendering at 2K in durations from 5 to 15 seconds across seven aspect ratios.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-video](/content/models/minimax/h3/text-to-video/index.html)

[**Model Labs**](/content/explore/labs/index.html)

Explore the AI labs powering models on fal

\\
\\
Kling](/content/explore/kling/index.html) \\
\\
Minimax](/content/explore/minimax/index.html) \\
\\
LTX](/content/explore/ltx/index.html) \\
\\
xAI](/content/explore/xai/index.html) \\
\\
OpenAI](/content/explore/openai/index.html) \\
\\
Krea](/content/explore/krea/index.html) \\
\\
ElevenLabs](/content/explore/elevenlabs/index.html) \\
\\
Bytedance](/content/explore/bytedance/index.html) \\
\\
Alibaba](/content/explore/alibaba/index.html) \\
\\
Google](/content/explore/google/index.html) \\
\\
Bria AI](/content/explore/bria-ai/index.html) \\
\\
Black Forest Labs](/content/explore/black-forest-labs/index.html) \\
\\
Veed](/content/explore/veed/index.html) \\
\\
Heygen](/content/explore/heygen/index.html) \\
\\
Luma AI](/content/explore/luma-ai/index.html) \\
\\
Ideogram](/content/explore/ideogram/index.html) \\
\\
Pixverse](/content/explore/pixverse/index.html)

[**Seedance 2.5**](/content/explore/seedance-2.5)

[**New and Noteworthy**](/content/explore/new-and-noteworthy/index.html)

State-of-the-art models we think you'll love!

\\
\\
new\\
\\
\\
alibaba/qwen-image-3/text-to-image\\
\\
Edits images from one to three reference images and a natural-language instruction, preserving key details such as facial features and identity while applying the requested changes\\
\\
stylized\\
\\
transform\\
\\
typography\\
\\
text-to-image](/content/models/alibaba/qwen-image-3/text-to-image/index.html)

\\
\\
\\
xai/grok-imagine-video/v1.5/image-to-video\\
\\
Generate videos from images with audio using xAI's Grok Imagine 1.5 Video model.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/xai/grok-imagine-video/v1.5/image-to-video/index.html)

\\
\\
\\
bytedance/seedream/v5/lite/edit\\
\\
Image editing endpoint for the fast Lite version of Seedream 5.0, supporting high quality intelligent image editing with multiple inputs.\\
\\
bytedance\\
\\
seedream-5.0-lite\\
\\
edit\\
\\
image-to-image](/content/models/fal-ai/bytedance/seedream/v5/lite/edit/index.html)

[**Grok Imagine**](/content/explore/grok-imagine/index.html)

\\
\\
\\
xai/grok-imagine-image/quality/edit\\
\\
Grok Imagine Pro is an advanced AI model from xAI that creates high-quality visuals from text prompts and allows you to edit or analyze existing images.\\
\\
stylized\\
\\
transform\\
\\
typography\\
\\
image-to-image](/content/models/xai/grok-imagine-image/quality/edit/index.html)

\\
\\
\\
xai/grok-imagine-image/quality/text-to-image\\
\\
Grok Imagine Pro is an advanced AI model from xAI that creates high-quality visuals from text prompts and allows you to edit or analyze existing images.\\
\\
stylized\\
\\
transform\\
\\
typography\\
\\
text-to-image](/content/models/xai/grok-imagine-image/quality/text-to-image/index.html)

\\
\\
\\
xai/grok-imagine-video/extend-video\\
\\
Extend videos with xAI's Grok Imagine video model\\
\\
video-edit\\
\\
v2v\\
\\
grok\\
\\
video-to-video](/content/models/xai/grok-imagine-video/extend-video/index.html)

\\
\\
\\
xai/tts/v1\\
\\
Generate speech with expressive and realistic voices from xAI\\
\\
text-to-speech](/content/models/xai/tts/v1/index.html)

\\
\\
\\
xai/grok-imagine-video/edit-video\\
\\
Edit videos using xAI's Grok Imagine\\
\\
video-edit\\
\\
v2v\\
\\
grok\\
\\
video-to-video](/content/models/xai/grok-imagine-video/edit-video/index.html)

\\
\\
\\
xai/grok-imagine-video/reference-to-video\\
\\
Generate videos using multiple reference images with xAI's Grok Imagine video model\\
\\
video-edit\\
\\
v2v\\
\\
grok\\
\\
image-to-video](/content/models/xai/grok-imagine-video/reference-to-video/index.html)

\\
\\
\\
xai/grok-imagine-video/text-to-video\\
\\
Generate videos with audio from text using Grok Imagine Video.\\
\\
xai\\
\\
grok\\
\\
t2v\\
\\
text-to-video](/content/models/xai/grok-imagine-video/text-to-video/index.html)

\\
\\
\\
xai/grok-imagine-video/image-to-video\\
\\
Generate videos from images with audio using xAI's Grok Imagine Video model. \\
\\
grok\\
\\
xai\\
\\
i2v\\
\\
image-to-video](/content/models/xai/grok-imagine-video/image-to-video/index.html)

\\
\\
\\
xai/grok-imagine-image\\
\\
Generate highly aesthetic images with xAI's Grok Imagine Image generation model.\\
\\
xai\\
\\
grok\\
\\
text-to-image](/content/models/xai/grok-imagine-image/index.html)

[**Best AI Image Generators**](/content/explore/best-ai-image-generators/index.html)

Unlock the future of creativity with these text to image, AI image generator models.

[![Super fast endpoint for the FLUX.1 [dev] model with LoRA support, enabling rapid and high-quality image generation using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a9f91af%2FUMrx6t6mc59z33d20WIQP_iBIZYKKq.png/tr:w-1920,q-80/UMrx6t6mc59z33d20WIQP_iBIZYKKq.webp)\\
\\
\\
flux-lora\\
\\
Super fast endpoint for the FLUX.1 \[dev\] model with LoRA support, enabling rapid and high-quality image generation using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.\\
\\
lora\\
\\
personalization\\
\\
text-to-image](/content/models/fal-ai/flux-lora/index.html)

\\
\\
\\
google/nano-banana-2-lite\\
\\
Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.\\
\\
text-to-image](/content/models/google/nano-banana-2-lite/index.html)

\\
\\
\\
z-image/turbo\\
\\
Z-Image Turbo is a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.\\
\\
turbo\\
\\
z-image\\
\\
fast\\
\\
text-to-image](/content/models/fal-ai/z-image/turbo/index.html)

\\
\\
\\
bytedance/seedream/v4.5/text-to-image\\
\\
A new-generation image creation model ByteDance, Seedream 4.5 integrates image generation and image editing capabilities into a single, unified architecture.\\
\\
stylized\\
\\
transform\\
\\
text-to-image](/content/models/fal-ai/bytedance/seedream/v4.5/text-to-image/index.html)

\\
\\
\\
bytedance/seedream/v4/text-to-image\\
\\
A new-generation image creation model ByteDance, Seedream 4.0 integrates image generation and image editing capabilities into a single, unified architecture.\\
\\
stylized\\
\\
transform\\
\\
text-to-image](/content/models/fal-ai/bytedance/seedream/v4/text-to-image/index.html)

[![Text-to-image generation with FLUX.2 [dev] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2Fpenguin%2FzSBCJtPpeIQwR5AC_IamX_b1e1137961754e4d851907c21f8c20cd.jpg/tr:w-1920,q-80/zSBCJtPpeIQwR5AC_IamX_b1e1137961754e4d851907c21f8c20cd.webp)\\
\\
\\
flux-2\\
\\
Text-to-image generation with FLUX.2 \[dev\] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.\\
\\
text-to-image](/content/models/fal-ai/flux-2/index.html)

\\
\\
\\
ideogram/v3\\
\\
Generate high-quality images, posters, and logos with Ideogram V3. Features exceptional typography handling and realistic outputs optimized for commercial and creative use.\\
\\
realism\\
\\
typography\\
\\
text-to-image](/content/models/fal-ai/ideogram/v3/index.html)

\\
\\
\\
krea/v2/large/text-to-image\\
\\
Generate high-fidelity images from text with Krea 2 Large, supporting aspect ratio, creativity, seed controls, and optional style references.\\
\\
image-generation\\
\\
style-reference\\
\\
krea\\
\\
text-to-image](/content/models/krea/v2/large/text-to-image/index.html)

\\
\\
recraft/v3/text-to-image\\
\\
Recraft V3 is a text-to-image model with the ability to generate long texts, vector art, images in brand style, and much more. As of today, it is SOTA in image generation, proven by Hugging Face's industry-leading Text-to-Image Benchmark by Artificial Analysis.\\
\\
vector\\
\\
typography\\
\\
style\\
\\
text-to-image](/content/models/fal-ai/recraft/v3/text-to-image/index.html)

\\
\\
\\
gpt-image-1.5\\
\\
GPT Image 1.5 generates high-fidelity images with strong prompt adherence, preserving composition, lighting, and fine-grained detail.\\
\\
openai\\
\\
gpt-image\\
\\
text-to-image](/content/models/fal-ai/gpt-image-1.5)

[![Text-to-image generation with FLUX.2 [dev] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities—all at turbo speed.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a871494%2Fj8F-tmy_dz4TyImvIHj19_510cc93373ef451386734b7e05711de1.jpg/tr:w-1920,q-80/j8F-tmy_dz4TyImvIHj19_510cc93373ef451386734b7e05711de1.webp)\\
\\
\\
flux-2/turbo\\
\\
Text-to-image generation with FLUX.2 \[dev\] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities—all at turbo speed.\\
\\
text-to-image](/content/models/fal-ai/flux-2/turbo/index.html)

\\
\\
fast-sdxl\\
\\
Run SDXL at the speed of light\\
\\
diffusion\\
\\
lora\\
\\
embeddings\\
\\
text-to-image](/content/models/fal-ai/fast-sdxl/index.html)

[![Text-to-image generation with FLUX.2 [dev] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities— in a flash.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a871486%2FtX7YdfQViGtCE7ZjxOCph_5f5262a21e9e426e8981ea9513d11999.jpg/tr:w-1920,q-80/tX7YdfQViGtCE7ZjxOCph_5f5262a21e9e426e8981ea9513d11999.webp)\\
\\
\\
flux-2/flash\\
\\
Text-to-image generation with FLUX.2 \[dev\] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities— in a flash.\\
\\
text-to-image](/content/models/fal-ai/flux-2/flash/index.html)

\\
\\
\\
gemini-3-pro-image-preview\\
\\
Gemini 3 Pro Image (a.k.a Nano Banana Pro) is Google's state-of-the-art high-fidelity image generation and editing model\\
\\
realism\\
\\
typography\\
\\
text-to-image](/content/models/fal-ai/gemini-3-pro-image-preview/index.html)

\\
\\
\\
gemini-25-flash-image\\
\\
Google's famous original image generation and editing model, a.k.a Nano Banana\\
\\
text-to-image](/content/models/fal-ai/gemini-25-flash-image/index.html)

[![Fastest inference in the world for the 12 billion parameter FLUX.1 [schnell] text-to-image model. ](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffal_cdn%2Ffal%2FUpscale-2.jpg/tr:w-1920,q-80/Upscale-2.webp)\\
\\
\\
flux-1/schnell\\
\\
Fastest inference in the world for the 12 billion parameter FLUX.1 \[schnell\] text-to-image model. \\
\\
text-to-image](/content/models/fal-ai/flux-1/schnell/index.html)

[**Best Image Editing Models**](/content/explore/best-image-editing-models/index.html)

The fan favorite best image editing models on the market

\\
\\
reve/edit\\
\\
Reve’s edit model lets you upload an existing image and then transform it via a text prompt\\
\\
image-to-image](/content/models/fal-ai/reve/edit/index.html)

\\
\\
\\
bria/fibo-edit/edit\\
\\
High-fidelity image editing model with state-of-the-art controllability. Combines JSON + Mask + Image for precise, fine-grained edits ideal for production and enterprise workflows. Trained on licensed data - safe for commercial use.\\
\\
bria\\
\\
fibo-edit\\
\\
image-editing\\
\\
image-to-image](/content/models/bria/fibo-edit/edit/index.html)

\\
\\
\\
bytedance/seedream/v4/edit\\
\\
A new-generation image creation model ByteDance, Seedream 4.0 integrates image generation and image editing capabilities into a single, unified architecture.\\
\\
stylized\\
\\
transform\\
\\
editing\\
\\
image-to-image](/content/models/fal-ai/bytedance/seedream/v4/edit/index.html)

[![Fast endpoint for the FLUX.1 Kontext [dev] model with LoRA support, enabling rapid and high-quality image editing using pre-trained LoRA adaptations for specific styles, brand identities, and product-specific outputs.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffal_cdn%2Ffal%2FUpscale-3.jpeg/tr:w-1920,q-80/Upscale-3.webp)\\
\\
\\
flux-kontext-lora\\
\\
Fast endpoint for the FLUX.1 Kontext \[dev\] model with LoRA support, enabling rapid and high-quality image editing using pre-trained LoRA adaptations for specific styles, brand identities, and product-specific outputs.\\
\\
image-editing\\
\\
image-to-image](/content/models/fal-ai/flux-kontext-lora/index.html)

\\
\\
\\
bria/fibo/generate\\
\\
SOTA open-source text-to-image model delivering high-fidelity outputs with accurate typography. JSON-structured prompts provide production-ready controllability for enterprise and agentic workflows. Trained exclusively on licensed data.\\
\\
bria\\
\\
fibo\\
\\
prompt-adherence\\
\\
text-to-image](/content/models/bria/fibo/generate/index.html)

[**Best of Open Source**](/content/explore/best-of-open-source/index.html)

Some of our favorite open source media models

[![LoRA trainer for FLUX.1 Kontext [dev]](https://refinery.fal.media/url/https%3A%2F%2Fv3.fal.media%2Ffiles%2Fmonkey%2FpYXiffttc2Skv36wflufu_dec4efe0d27e4527b64acfbc0e91536a.jpg/tr:w-1920,q-80/pYXiffttc2Skv36wflufu_dec4efe0d27e4527b64acfbc0e91536a.webp)\\
\\
\\
flux-kontext-trainer\\
\\
LoRA trainer for FLUX.1 Kontext \[dev\]\\
\\
training](/content/models/fal-ai/flux-kontext-trainer/index.html)

\\
\\
\\
ltx-video-13b-distilled/image-to-video\\
\\
Generate videos from prompts and images using LTX Video-0.9.7 13B Distilled and custom LoRA\\
\\
video\\
\\
ltx-video\\
\\
image-to-video](/content/models/fal-ai/ltx-video-13b-distilled/image-to-video/index.html)

\\
\\
\\
wan-22-image-trainer\\
\\
Wan 2.2 text to image LoRA trainer. Fine-tune Wan 2.2 for subjects and styles with unprecedented detail.\\
\\
lora\\
\\
personalization\\
\\
training](/content/models/fal-ai/wan-22-image-trainer/index.html)

\\
\\
\\
wan/v2.2-a14b/image-to-video/lora\\
\\
Wan-2.2 image-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts and images. This endpoint supports LoRAs made for Wan 2.2\\
\\
motion\\
\\
lora\\
\\
image-to-video](/content/models/fal-ai/wan/v2.2-a14b/image-to-video/lora/index.html)

\\
\\
\\
qwen-image\\
\\
Qwen-Image is an image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. \\
\\
text-to-image](/content/models/fal-ai/qwen-image/index.html)

[![Super fast endpoint for the FLUX.1 [dev] model with LoRA support, enabling rapid and high-quality image generation using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a9f9a61%2FT4z71gOSeWv0wALDdi2-b_qVoN8eec.png/tr:w-1920,q-80/T4z71gOSeWv0wALDdi2-b_qVoN8eec.webp)\\
\\
\\
flux-krea-lora/stream\\
\\
Super fast endpoint for the FLUX.1 \[dev\] model with LoRA support, enabling rapid and high-quality image generation using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.\\
\\
lora\\
\\
personalization\\
\\
text-to-image](/content/models/fal-ai/flux-krea-lora/stream/index.html)

[**Seedance 2.0**](/content/explore/seedance-2.0)

The new sota video model by Bytedance. Access the new stunning video generation model today.

\\
\\
\\
bytedance/seedance-2.0/text-to-video\\
\\
ByteDance's most advanced text-to-video model. Cinematic output with native audio, multi-shot editing, real-world physics, and director-level camera control.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-video](/content/models/bytedance/seedance-2.0/text-to-video/index.html)

\\
\\
\\
bytedance/seedance-2.0/fast/text-to-video\\
\\
ByteDance's most advanced text-to-video model, fast tier. Lower latency and cost with cinematic output, native audio, multi-shot editing, and director-level camera control.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-video](/content/models/bytedance/seedance-2.0/fast/text-to-video/index.html)

\\
\\
\\
bytedance/seedance-2.0/fast/reference-to-video\\
\\
ByteDance's most advanced reference-to-video model, fast tier. Lower latency and cost with up to 9 images, 3 videos, and 3 audio clips as inputs.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/bytedance/seedance-2.0/fast/reference-to-video/index.html)

\\
\\
\\
bytedance/seedance-2.0/fast/image-to-video\\
\\
ByteDance's most advanced image-to-video model, fast tier. Lower latency and cost with synchronized audio, start and end frame control, and motion prompts.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/bytedance/seedance-2.0/fast/image-to-video/index.html)

[**Text To Speech APIs**](/content/explore/text-to-speech-apis/index.html)

Create lifelike speech with our AI text to speech APIs

\\
\\
\\
gemini-3.1-flash-tts\\
\\
Newest audio model from Google introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.\\
\\
lipsync\\
\\
avatar\\
\\
text-to-speech](/content/models/fal-ai/gemini-3.1-flash-tts/index.html)

\\
\\
\\
minimax/speech-2.8-hd\\
\\
Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.\\
\\
text-to-speech](/content/models/fal-ai/minimax/speech-2.8-hd/index.html)

\\
\\
\\
minimax/speech-2.8-turbo\\
\\
Generate speech from text prompts and different voices using the MiniMax Speech-2.8 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.\\
\\
text-to-speech](/content/models/fal-ai/minimax/speech-2.8-turbo/index.html)

\\
\\
\\
xai/tts/v1\\
\\
Generate speech with expressive and realistic voices from xAI\\
\\
text-to-speech](/content/models/xai/tts/v1/index.html)

\\
\\
\\
qwen-3-tts/text-to-speech/1.7b\\
\\
Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model\\
\\
text-to-speech](/content/models/fal-ai/qwen-3-tts/text-to-speech/1.7b)

\\
\\
inworld-tts\\
\\
Text to Speech Endpoint for Inworld's TTS-1.5 Max.\\
\\
inworld\\
\\
tts\\
\\
text-to-speech](/content/models/fal-ai/inworld-tts/index.html)

\\
\\
lux-tts\\
\\
High-quality voice cloning TTS model that generates 48kHz speech from text and a reference audio. Distilled to 4 steps for fast inference.\\
\\
deprecated\\
\\
tts\\
\\
voice-cloning\\
\\
speech-synthesis\\
\\
text-to-speech](/content/models/fal-ai/lux-tts/index.html)

\\
\\
\\
elevenlabs/tts/turbo-v2.5\\
\\
Generate high-speed text-to-speech audio using ElevenLabs TTS Turbo v2.5.\\
\\
audio\\
\\
text-to-speech](/content/models/fal-ai/elevenlabs/tts/turbo-v2.5)

\\
\\
chatterbox/text-to-speech\\
\\
Whether you're working on memes, videos, games, or AI agents, Chatterbox brings your content to life. Use the first tts from resemble ai.\\
\\
text-to-speech](/content/models/fal-ai/chatterbox/text-to-speech/index.html)

\\
\\
maya/batch\\
\\
Maya1 is a state-of-the-art speech model by Maya Research for expressive voice generation, built to capture real human emotion and precise voice design.\\
\\
tts\\
\\
text-to-speech](/content/models/fal-ai/maya/batch/index.html)

\\
\\
\\
minimax/speech-02-hd\\
\\
Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.\\
\\
speech\\
\\
text-to-speech](/content/models/fal-ai/minimax/speech-02-hd/index.html)

\\
\\
\\
minimax/voice-clone\\
\\
Clone a voice from a sample audio and generate speech from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality text-to-speech.\\
\\
speech\\
\\
text-to-speech](/content/models/fal-ai/minimax/voice-clone/index.html)

\\
\\
\\
minimax/speech-02-turbo\\
\\
Generate fast speech from text prompts and different voices using the MiniMax Speech-02 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.\\
\\
speech\\
\\
text-to-speech](/content/models/fal-ai/minimax/speech-02-turbo/index.html)

\\
\\
\\
minimax/speech-2.6-hd\\
\\
Generate speech from text prompts and different voices using the MiniMax Speech-2.6 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.\\
\\
text-to-speech](/content/models/fal-ai/minimax/speech-2.6-hd/index.html)

\\
\\
index-tts-2/text-to-speech\\
\\
Generate natural, clear speeches using Index TTS 2.0 from IndexTeam\\
\\
text-to-speech](/content/models/fal-ai/index-tts-2/text-to-speech/index.html)

\\
\\
\\
minimax/speech-2.6-turbo\\
\\
Generate speech from text prompts and different voices using the MiniMax Speech-2.6 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.\\
\\
text-to-speech](/content/models/fal-ai/minimax/speech-2.6-turbo/index.html)

\\
\\
vibevoice/7b\\
\\
Generate long, expressive multi-voice speech using Microsoft's powerful TTS\\
\\
multi-speaker\\
\\
podcast\\
\\
text-to-speech](/content/models/fal-ai/vibevoice/7b/index.html)

\\
\\
\\
kling-video/v1/tts\\
\\
Generate speech from text prompts and different voices using the Kling TTS model, which leverages advanced AI techniques to create high-quality text-to-speech.\\
\\
audio\\
\\
text-to-speech](/content/models/fal-ai/kling-video/v1/tts/index.html)

\\
\\
\\
qwen-3-tts/voice-design/1.7b\\
\\
Create custom voices using Qwen3-TTS Voice Design model and later use Clone Voice model to create your own voices!\\
\\
voice-design\\
\\
text-to-speech](/content/models/fal-ai/qwen-3-tts/voice-design/1.7b)

\\
\\
\\
qwen-3-tts/text-to-speech/0.6b\\
\\
Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model\\
\\
text-to-speech](/content/models/fal-ai/qwen-3-tts/text-to-speech/0.6b)

[**AI Image Generator APIs**](/content/explore/ai-image-generator-apis/index.html)

Generate a variety of stunning images using our AI Image Generator APIs

\\
\\
fast-sdxl\\
\\
Run SDXL at the speed of light\\
\\
diffusion\\
\\
lora\\
\\
embeddings\\
\\
text-to-image](/content/models/fal-ai/fast-sdxl/index.html)

[**Text to Video APIs**](/content/explore/text-to-video-apis/index.html)

Access the top Text to Video APIs with lightning fast inference speeds

\\
\\
\\
kling-video/v3/pro/text-to-video\\
\\
Kling 3.0 Pro: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation, with multi-shot support.\\
\\
text-to-video](/content/models/fal-ai/kling-video/v3/pro/text-to-video/index.html)

\\
\\
\\
kling-video/v3/standard/text-to-video\\
\\
Kling 3.0 Standard: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation, with multi-shot support.\\
\\
text-to-video](/content/models/fal-ai/kling-video/v3/standard/text-to-video/index.html)

\\
\\
\\
veo3.1\\
\\
Veo 3.1 by Google, the most advanced AI video generation model in the world. With sound on!\\
\\
text-to-video](/content/models/fal-ai/veo3.1)

\\
\\
\\
veo3.1/fast\\
\\
Faster and more cost effective version of Google's Veo 3.1! \\
\\
text-to-video](/content/models/fal-ai/veo3.1/fast/index.html)

\\
\\
\\
kling-video/v2.5-turbo/pro/text-to-video\\
\\
Kling 2.5 Turbo Pro: Top-tier text-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.\\
\\
animation\\
\\
stylized\\
\\
text-to-video](/content/models/fal-ai/kling-video/v2.5-turbo/pro/text-to-video/index.html)

\\
\\
\\
google/gemini-omni-flash\\
\\
Creates video with synchronized audio from text input. Grounded in Gemini's real-world knowledge, with improved physics understanding for more coherent motion and interaction.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-video](/content/models/google/gemini-omni-flash/index.html)

\\
\\
\\
bytedance/seedance/v1.5/pro/text-to-video\\
\\
Generate videos with audio with Seedance 1.5\\
\\
bytedance\\
\\
seedance\\
\\
audio\\
\\
text-to-video](/content/models/fal-ai/bytedance/seedance/v1.5/pro/text-to-video/index.html)

\\
\\
\\
kling-video/v2.6/pro/text-to-video\\
\\
Kling 2.6 Pro: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation.\\
\\
text-to-video](/content/models/fal-ai/kling-video/v2.6/pro/text-to-video/index.html)

\\
\\
\\
kling-video/lipsync/audio-to-video\\
\\
Kling LipSync is an audio-to-video model that generates realistic lip movements from audio input.\\
\\
audio to video\\
\\
lipsync\\
\\
text-to-video](/content/models/fal-ai/kling-video/lipsync/audio-to-video/index.html)

\\
\\
\\
veo3.1/lite\\
\\
Veo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-video](/content/models/fal-ai/veo3.1/lite/index.html)

\\
\\
\\
kling-video/v1.6/standard/text-to-video\\
\\
Generate video clips from your prompts using Kling 1.6 (std)\\
\\
text-to-video](/content/models/fal-ai/kling-video/v1.6/standard/text-to-video/index.html)

\\
\\
\\
bytedance/seedance-2.0/mini/text-to-video\\
\\
Seedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-video](/content/models/bytedance/seedance-2.0/mini/text-to-video/index.html)

\\
\\
\\
kling-video/o3/pro/text-to-video\\
\\
Generate realistic videos using Kling O3 from Kling Team!\\
\\
text-to-video](/content/models/fal-ai/kling-video/o3/pro/text-to-video/index.html)

\\
\\
\\
bytedance/seedance/v1/pro/text-to-video\\
\\
Seedance 1.0 Pro, a high quality video generation model developed by Bytedance.\\
\\
text-to-video](/content/models/fal-ai/bytedance/seedance/v1/pro/text-to-video/index.html)

\\
\\
\\
wan/v2.7/text-to-video\\
\\
Wan 2.7 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-video](/content/models/fal-ai/wan/v2.7/text-to-video/index.html)

\\
\\
\\
pixverse/v6/text-to-video\\
\\
Pixverse's latest v6 Model.\\
\\
text-to-video](/content/models/fal-ai/pixverse/v6/text-to-video/index.html)

\\
\\
\\
kling-video/o3/standard/text-to-video\\
\\
Generate realistic videos using Kling O3 from Kling Team!\\
\\
text-to-video](/content/models/fal-ai/kling-video/o3/standard/text-to-video/index.html)

\\
\\
\\
minimax/hailuo-02/standard/text-to-video\\
\\
MiniMax Hailuo-02 Text To Video API (Standard, 768p): Advanced video generation model with 768p resolution\\
\\
text-to-video](/content/models/fal-ai/minimax/hailuo-02/standard/text-to-video/index.html)

\\
\\
\\
wan-25-preview/text-to-video\\
\\
Wan 2.5 text-to-video model.\\
\\
text-to-video](/content/models/fal-ai/wan-25-preview/text-to-video/index.html)

\\
\\
\\
kling-video/v3/turbo/standard/text-to-video\\
\\
Kling 3.0 Turbo Standard is a fast, cost-efficient video generation model that turns text prompts directly into 720P video with native audio, optimized for rapid iteration and high-volume production\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-video](/content/models/fal-ai/kling-video/v3/turbo/standard/text-to-video/index.html)

\\
\\
\\
ltx-2.3/text-to-video/fast\\
\\
LTX-2.3 is a high-quality, fast AI video model available in Pro and Fast variants for text-to-video, image-to-video, and audio-to-video.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-video](/content/models/fal-ai/ltx-2.3/text-to-video/fast/index.html)

\\
\\
\\
wan/v2.2-a14b/text-to-video\\
\\
Wan-2.2 text-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts. \\
\\
text to video\\
\\
motion\\
\\
text-to-video](/content/models/fal-ai/wan/v2.2-a14b/text-to-video/index.html)

[**Image to Video APIs**](/content/explore/image-to-video-apis/index.html)

\\
\\
\\
kling-video/v2.6/pro/image-to-video\\
\\
Kling 2.6 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation.\\
\\
image-to-video](/content/models/fal-ai/kling-video/v2.6/pro/image-to-video/index.html)

\\
\\
new\\
\\
\\
minimax/h3/reference-to-video\\
\\
MiniMax H3 is a frontier video model. This endpoint generates 2K video from multimodal references up to 9 images for subject and style, 3 video clips for motion, and 3 audio clips each cited in the prompt by order, keeping subjects consistent while following the referenced motion and audio.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/minimax/h3/reference-to-video/index.html)

\\
\\
\\
veo3.1/fast/image-to-video\\
\\
Generate videos from your image prompts using Veo 3.1 fast.\\
\\
image-to-video](/content/models/fal-ai/veo3.1/fast/image-to-video/index.html)

\\
\\
\\
bytedance/seedance/v1.5/pro/image-to-video\\
\\
Generate videos with audio with Seedance 1.5 (supports start & end frame) \\
\\
bytedance\\
\\
seedance\\
\\
audio\\
\\
image-to-video](/content/models/fal-ai/bytedance/seedance/v1.5/pro/image-to-video/index.html)

\\
\\
\\
veo3.1/image-to-video\\
\\
Veo 3.1 is the latest state-of-the art video generation model from Google DeepMind\\
\\
image-to-video](/content/models/fal-ai/veo3.1/image-to-video/index.html)

\\
\\
\\
kling-video/v2.1/standard/image-to-video\\
\\
Kling 2.1 Standard is a cost-efficient endpoint for the Kling 2.1 model, delivering high-quality image-to-video generation \\
\\
image-to-video](/content/models/fal-ai/kling-video/v2.1/standard/image-to-video/index.html)

\\
\\
\\
google/gemini-omni-flash/image-to-video\\
\\
Animates a still image into video with audio. Extends a single frame into coherent motion, grounded in Gemini's physical understanding of how scenes and subjects behave.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/google/gemini-omni-flash/image-to-video/index.html)

\\
\\
\\
veo3.1/lite/image-to-video\\
\\
Veo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/fal-ai/veo3.1/lite/image-to-video/index.html)

\\
\\
\\
kling-video/o3/pro/image-to-video\\
\\
Generate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.\\
\\
image-to-video](/content/models/fal-ai/kling-video/o3/pro/image-to-video/index.html)

\\
\\
\\
bytedance/seedance/v1/pro/image-to-video\\
\\
Seedance 1.0 Pro, a high quality video generation model developed by Bytedance.\\
\\
image-to-video](/content/models/fal-ai/bytedance/seedance/v1/pro/image-to-video/index.html)

\\
\\
\\
kling-video/o3/standard/image-to-video\\
\\
Generate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.\\
\\
image-to-video](/content/models/fal-ai/kling-video/o3/standard/image-to-video/index.html)

\\
\\
\\
kling-video/o3/pro/reference-to-video\\
\\
Transform images, elements, and text into consistent, high-quality video scenes, ensuring stable character identity, object details, and environments.\\
\\
reference-to-video\\
\\
image-to-video](/content/models/fal-ai/kling-video/o3/pro/reference-to-video/index.html)

\\
\\
\\
google/gemini-omni-flash/reference-to-video\\
\\
Generates video with audio from combined multimodal references. Accepts text, images, audio, and video together as input to guide subject, motion, style, and sound in the output.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/google/gemini-omni-flash/reference-to-video/index.html)

\\
\\
\\
wan/v2.7/image-to-video\\
\\
Wan 2.7 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/fal-ai/wan/v2.7/image-to-video/index.html)

\\
\\
\\
kling-video/v1.6/standard/image-to-video\\
\\
Generate video clips from your images using Kling 1.6 (std)\\
\\
image-to-video](/content/models/fal-ai/kling-video/v1.6/standard/image-to-video/index.html)

\\
\\
\\
bytedance/seedance-2.0/mini/reference-to-video\\
\\
Seedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
image-to-video](/content/models/bytedance/seedance-2.0/mini/reference-to-video/index.html)

\\
\\
\\
minimax/hailuo-02/standard/image-to-video\\
\\
MiniMax Hailuo-02 Image To Video API (Standard, 768p, 512p): Advanced image-to-video generation model with 768p and 512p resolutions\\
\\
image-to-video](/content/models/fal-ai/minimax/hailuo-02/standard/image-to-video/index.html)

[**Best Image Models**](/content/explore/best-image-models/index.html)

Top-performing models for high-quality image generation and editing.

\\
\\
recraft/v4/pro/text-to-image\\
\\
Recraft V4 was developed with designers to bring true visual taste to AI image generation. Built for brand systems and production-ready workflows, it goes beyond prompt accuracy — delivering stronger composition, refined lighting, realistic materials, and a cohesive aesthetic. The result is imagery shaped by professional design judgment, ready for immediate real-world use without additional post-processing.\\
\\
text-to-image](/content/models/fal-ai/recraft/v4/pro/text-to-image/index.html)

\\
\\
imagineart/imagineart-2.0-preview/text-to-image\\
\\
ImagineArt 2.0 is ImagineArt's latest state-of-the-art visual reasoning text-to-image model, generating high-fidelity, professional-grade visuals with lifelike realism, cinematic effects, and strong aesthetic quality.\\
\\
stylized\\
\\
transform\\
\\
typography\\
\\
text-to-image](/content/models/imagineart/imagineart-2.0-preview/text-to-image/index.html)

[**Background Remover APIs**](/content/explore/background-remover-apis/index.html)

Find the API of your choice to remove a background from your image or video

\\
\\
pixelcut/background-removal\\
\\
Pixelcut’s Background Remover enables fast, ultra high-quality removal of backgrounds from images. Perfect for e-commerce and image editing workflows. Powered by advanced AI for clean, perfect cutouts every time.\\
\\
background removal\\
\\
utility\\
\\
remove background\\
\\
image-to-image](/content/models/pixelcut/background-removal/index.html)

\\
\\
\\
ideogram/remove-background\\
\\
Remove backgrounds from existing images with Ideogram's remove background feature. Isolate subjects cleanly for compositing and creative reuse.\\
\\
image-to-image](/content/models/fal-ai/ideogram/remove-background/index.html)

\\
\\
\\
bria/video/background-removal/v3\\
\\
Remove backgrounds from any video with Bria's VRMBG 3.0. Fast, accurate background removal across talking heads, podcasts, product videos, commercials, and cinematic footage.\\
\\
video-to-video](/content/models/bria/video/background-removal/v3/index.html)

\\
\\
imageutils/rembg\\
\\
Remove the background from an image.\\
\\
background removal\\
\\
utility\\
\\
editing\\
\\
image-to-image](/content/models/fal-ai/imageutils/rembg/index.html)

\\
\\
birefnet\\
\\
bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)\\
\\
background removal\\
\\
segmentation\\
\\
high-res\\
\\
image-to-image](/content/models/fal-ai/birefnet/index.html)

\\
\\
\\
veed/video-background-removal/green-screen\\
\\
Remove background from videos filmed using chromakey, with automatic green spill suppression for clean, professional edges.\\
\\
video-to-video](/content/models/veed/video-background-removal/green-screen/index.html)

\\
\\
\\
veed/video-background-removal/fast\\
\\
Remove background from any video with people and objects. No green screen needed.\\
\\
video-to-video](/content/models/veed/video-background-removal/fast/index.html)

\\
\\
\\
veed/video-background-removal\\
\\
Remove background from any video with people and objects. No green screen needed.\\
\\
video-to-video](/content/models/veed/video-background-removal/index.html)

[**Veo 3.1**](/content/explore/veo-3.1)

\\
\\
\\
veo3.1/fast/first-last-frame-to-video\\
\\
Generate videos from a first/last frame using Google's Veo 3.1 Fast\\
\\
image-to-video](/content/models/fal-ai/veo3.1/fast/first-last-frame-to-video/index.html)

\\
\\
\\
veo3.1/first-last-frame-to-video\\
\\
Generate videos from a first and last framed using Google's Veo 3.1\\
\\
image-to-video](/content/models/fal-ai/veo3.1/first-last-frame-to-video/index.html)

\\
\\
\\
veo3.1/fast\\
\\
Faster and more cost effective version of Google's Veo 3.1! \\
\\
text-to-video](/content/models/fal-ai/veo3.1/fast/index.html)

\\
\\
\\
veo3.1/reference-to-video\\
\\
Generate Videos from images using Google's Veo 3.1\\
\\
image-to-video](/content/models/fal-ai/veo3.1/reference-to-video/index.html)

[**Marquee Video Models**](/content/explore/marquee-video-models/index.html)

Flagship video generation models known for top-tier quality, motion control, and cinematic results.

\\
\\
\\
pixverse/v6/image-to-video\\
\\
Pixverse's latest V6 Model\\
\\
image-to-video](/content/models/fal-ai/pixverse/v6/image-to-video/index.html)

\\
\\
\\
kling-video/v2.1/pro/image-to-video\\
\\
Kling 2.1 Pro is an advanced endpoint for the Kling 2.1 model, offering professional-grade videos with enhanced visual fidelity, precise camera movements, and dynamic motion control, perfect for cinematic storytelling.\\
\\
image-to-video](/content/models/fal-ai/kling-video/v2.1/pro/image-to-video/index.html)

\\
\\
\\
wan/v2.2-a14b/image-to-video\\
\\
fal-ai/wan/v2.2-A14B/image-to-video\\
\\
image-to-video](/content/models/fal-ai/wan/v2.2-a14b/image-to-video/index.html)

\\
\\
\\
ltx-2-19b/image-to-video\\
\\
Generate video with audio from images using LTX-2\\
\\
image-to-video](/content/models/fal-ai/ltx-2-19b/image-to-video/index.html)

[**Best Avatar Models**](/content/explore/best-avatar-models/index.html)

Top models for generating talking avatars, lip-sync videos, and expressive character performances.

\\
\\
creatify/aurora\\
\\
Generate high fidelity, studio quality videos of your avatar speaking or singing using the Aurora from Creatify team!\\
\\
lipsync\\
\\
image-to-video](/content/models/fal-ai/creatify/aurora/index.html)

\\
\\
\\
veed/fabric-1.0\\
\\
VEED Fabric 1.0 is an image-to-video API that turns any image into a talking video\\
\\
lipsync\\
\\
avatar\\
\\
image-to-video](/content/models/veed/fabric-1.0)

\\
\\
\\
heygen/avatar5/digital-twin\\
\\
Create natural HeyGen Avatar V digital twin videos from text or audio, with lip-sync, optional backgrounds, captions, and MP4/WebM output.\\
\\
avatar\\
\\
digital-twin\\
\\
talking-avatar\\
\\
text-to-video](/content/models/fal-ai/heygen/avatar5/digital-twin/index.html)

\\
\\
sync-lipsync/v3\\
\\
sync-3 most powerful lipsync model yet, featuring native visual intelligence for professional-quality video.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
video-to-video](/content/models/fal-ai/sync-lipsync/v3/index.html)

\\
\\
\\
bytedance/omnihuman/v1.5\\
\\
Omnihuman v1.5 is a new and improved version of Omnihuman. It generates video using an image of a human figure paired with an audio file. It produces vivid, high-quality videos where the character’s emotions and movements maintain a strong correlation with the audio.\\
\\
lipsync\\
\\
image-to-video](/content/models/fal-ai/bytedance/omnihuman/v1.5)

\\
\\
ai-avatar/single-text\\
\\
MultiTalk model generates a talking avatar video from an image and text. Converts text to speech automatically, then generates the avatar speaking with lip-sync.\\
\\
stylized\\
\\
transform\\
\\
image-to-video](/content/models/fal-ai/ai-avatar/single-text/index.html)

\\
\\
sync-lipsync/v2\\
\\
Generate realistic lipsync animations from audio using advanced algorithms for high-quality synchronization with Sync Lipsync 2.0 model\\
\\
animation\\
\\
lip sync\\
\\
video-to-video](/content/models/fal-ai/sync-lipsync/v2/index.html)

\\
\\
\\
kling-video/v2.1/master/image-to-video\\
\\
Kling 2.1 Master: The premium endpoint for Kling 2.1, designed for top-tier image-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.\\
\\
image-to-video](/content/models/fal-ai/kling-video/v2.1/master/image-to-video/index.html)

\\
\\
\\
pixverse/lipsync\\
\\
Generate realistic lipsync animations from audio using advanced algorithms for high-quality synchronization with PixVerse Lipsync model\\
\\
animation\\
\\
lip sync\\
\\
video-to-video](/content/models/fal-ai/pixverse/lipsync/index.html)

\\
\\
\\
kling-video/v1/pro/ai-avatar\\
\\
Kling AI Avatar Pro: The premium endpoint for creating avatar videos with realistic humans, animals, cartoons, or stylized characters\\
\\
stylized\\
\\
transform\\
\\
image-to-video](/content/models/fal-ai/kling-video/v1/pro/ai-avatar/index.html)

[**Audio Models**](/content/explore/audio-models/index.html)

Models for speech, music, sound effects, and audio generation across a wide range of use cases.

\\
\\
sonilo/v1.1/text-to-music\\
\\
Generates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact duration.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-audio](/content/models/sonilo/v1.1/text-to-music/index.html)

\\
\\
new\\
\\
sonilo/v1.1/text-to-sound-effects\\
\\
Generates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact duration.\\
\\
sfx\\
\\
audio\\
\\
effects\\
\\
text-to-audio](/content/models/sonilo/v1.1/text-to-sound-effects/index.html)

\\
\\
new\\
\\
sonilo/v1.1/video-to-sound-effects\\
\\
Analyzes a video and generates synchronized, royalty-free sound effects timed to visible actions. Returns the generated sound-effects audio track for commercial use.\\
\\
sfx\\
\\
audio\\
\\
effects\\
\\
video-to-audio](/content/models/sonilo/v1.1/video-to-sound-effects/index.html)

\\
\\
inworld-tts\\
\\
Text to Speech Endpoint for Inworld's TTS-1.5 Max.\\
\\
inworld\\
\\
tts\\
\\
text-to-speech](/content/models/fal-ai/inworld-tts/index.html)

\\
\\
playai/tts/dialog\\
\\
Generate natural-sounding multi-speaker dialogues, and audio. Perfect for expressive outputs, storytelling, games, animations, and interactive media.\\
\\
deprecated\\
\\
audio\\
\\
text-to-audio](/content/models/fal-ai/playai/tts/dialog/index.html)

\\
\\
dia-tts/voice-clone\\
\\
Clone dialog voices from a sample audio and generate dialogs from text prompts using the Dia TTS which leverages advanced AI techniques to create high-quality text-to-speech.\\
\\
speech\\
\\
audio-to-audio](/content/models/fal-ai/dia-tts/voice-clone/index.html)

\\
\\
mirelo-ai/sfx-v1/video-to-audio\\
\\
Generate synced sounds for any video, and return the new sound track (like MMAudio)\\
\\
sfx\\
\\
video-to-audio](/content/models/mirelo-ai/sfx-v1/video-to-audio/index.html)

\\
\\
mirelo-ai/sfx-v1/video-to-video\\
\\
Generate synced sounds for any video, and return it with its new sound track (like MMAudio)\\
\\
sfx\\
\\
video-to-video](/content/models/mirelo-ai/sfx-v1/video-to-video/index.html)

\\
\\
beatoven/music-generation\\
\\
Generate royalty-free instrumental music from electronic, hip hop, and indie rock to cinematic and classical genres. Perfect for games, films, social content, podcasts, and more.\\
\\
deprecated\\
\\
speech\\
\\
audio\\
\\
music\\
\\
text-to-audio](/content/models/beatoven/music-generation/index.html)

\\
\\
beatoven/sound-effect-generation\\
\\
Create professional-grade sound effects from animal and vehicle to nature, sci-fi, and otherworldly sounds. Perfect for films, games, and digital content.\\
\\
sfx\\
\\
audio\\
\\
effects\\
\\
text-to-audio](/content/models/beatoven/sound-effect-generation/index.html)

\\
\\
\\
minimax/preview/speech-2.5-hd\\
\\
Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.\\
\\
speech\\
\\
text-to-speech](/content/models/fal-ai/minimax/preview/speech-2.5-hd/index.html)

\\
\\
chatterbox/text-to-speech/multilingual\\
\\
Whether you're working on memes, videos, games, or AI agents, Chatterbox brings your content to life. Use the first tts from resemble ai.\\
\\
multilingual\\
\\
text-to-speech](/content/models/fal-ai/chatterbox/text-to-speech/multilingual/index.html)

\\
\\
resemble-ai/chatterboxhd/text-to-speech\\
\\
Generate expressive, natural speech with Resemble AI's Chatterbox. Features unique emotion control, instant voice cloning from short audio, and built-in watermarking.\\
\\
text-to-speech](/content/models/resemble-ai/chatterboxhd/text-to-speech/index.html)

\\
\\
\\
minimax/preview/speech-2.5-turbo\\
\\
Generate fast speech from text prompts and different voices using the MiniMax Speech-02 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.\\
\\
text-to-speech](/content/models/fal-ai/minimax/preview/speech-2.5-turbo/index.html)

\\
\\
\\
minimax-music/v2\\
\\
Generate music from text prompts using the MiniMax Music 2.0 model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.\\
\\
music\\
\\
audio\\
\\
text-to-audio](/content/models/fal-ai/minimax-music/v2/index.html)

\\
\\
\\
elevenlabs/music\\
\\
Generate high quality, realistic music with fine controls using Elevenlabs Music!\\
\\
music\\
\\
text-to-music\\
\\
text-to-audio](/content/models/fal-ai/elevenlabs/music/index.html)

\\
\\
cassetteai/music-generator\\
\\
CassetteAI’s model generates a 30-second sample in under 2 seconds and a full 3-minute track in under 10 seconds. At 44.1 kHz stereo audio, expect a level of professional consistency with no breaks, no squeaks, and no random interruptions in your creations.\\
\\
music\\
\\
cassetteai\\
\\
text-to-audio](/content/models/cassetteai/music-generator/index.html)

\\
\\
\\
minimax-music/v2.6\\
\\
MiniMax Music 2.6 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-audio](/content/models/fal-ai/minimax-music/v2.6)

[**Text to Music APIs**](/content/explore/text-to-music-apis/index.html)

Everything you need to start making music with AI

\\
\\
\\
lyria3/pro\\
\\
Lyria 3 Pro is the latest music model from Google\\
\\
audio\\
\\
sfx\\
\\
text-to-audio](/content/models/fal-ai/lyria3/pro/index.html)

\\
\\
\\
lyria3\\
\\
Lyria 3 is most recent music model from Google\\
\\
audio\\
\\
music\\
\\
sfx\\
\\
text-to-audio](/content/models/fal-ai/lyria3/index.html)

\\
\\
\\
minimax-music/v2.5\\
\\
MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.\\
\\
stylized\\
\\
transform\\
\\
lipsync\\
\\
text-to-audio](/content/models/fal-ai/minimax-music/v2.5)

\\
\\
\\
minimax-music\\
\\
Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.\\
\\
music\\
\\
text-to-audio](/content/models/fal-ai/minimax-music/index.html)

\\
\\
\\
minimax-music/v1.5\\
\\
Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.\\
\\
music\\
\\
text-to-audio](/content/models/fal-ai/minimax-music/v1.5)

\\
\\
stable-audio-25/text-to-audio\\
\\
Generate high quality music and sound effects using Stable Audio 2.5 from StabilityAI\\
\\
audio\\
\\
text-to-audio](/content/models/fal-ai/stable-audio-25/text-to-audio/index.html)

\\
\\
ace-step\\
\\
Generate music with lyrics from text using ACE-Step\\
\\
text-to-music\\
\\
text-to-audio](/content/models/fal-ai/ace-step/index.html)

\\
\\
ace-step/prompt-to-audio\\
\\
Generate music from a simple prompt using ACE-Step\\
\\
text-to-music\\
\\
text-to-audio](/content/models/fal-ai/ace-step/prompt-to-audio/index.html)

\\
\\
yue\\
\\
YuE is a groundbreaking series of open-source foundation models designed for music generation, specifically for transforming lyrics into full songs.\\
\\
deprecated\\
\\
music\\
\\
text-to-audio](/content/models/fal-ai/yue/index.html)

\\
\\
stable-audio\\
\\
Open source text-to-audio model.\\
\\
music\\
\\
text-to-audio](/content/models/fal-ai/stable-audio/index.html)

[**Best Lora Trainers**](/content/explore/best-lora-trainers/index.html)

Training endpoints for creating and fine-tuning custom LoRA models for personalization and style adaptation.

\\
\\
\\
flux-lora-portrait-trainer\\
\\
FLUX LoRA training optimized for portrait generation, with bright highlights, excellent prompt following and highly detailed results.\\
\\
lora\\
\\
personalization\\
\\
training](/content/models/fal-ai/flux-lora-portrait-trainer/index.html)

\\
\\
\\
flux-lora-fast-training\\
\\
Train styles, people and other subjects at blazing speeds.\\
\\
lora\\
\\
personalization\\
\\
training](/content/models/fal-ai/flux-lora-fast-training/index.html)

\\
\\
\\
wan-trainer/t2v-14b\\
\\
Train custom LoRAs for Wan-2.1 T2V 14B\\
\\
lora\\
\\
training](/content/models/fal-ai/wan-trainer/t2v-14b/index.html)

\\
\\
\\
qwen-image-trainer\\
\\
Qwen Image LoRA training\\
\\
deprecated\\
\\
lora\\
\\
personalization\\
\\
training](/content/models/fal-ai/qwen-image-trainer/index.html)

[![Fine-tune FLUX.2 [klein] 4B from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a8b082e%2FN8Fy12FSedqMd-2Ehh8z1_70e50238ee6b479ebd61270840b4806e.jpg/tr:w-1920,q-80/N8Fy12FSedqMd-2Ehh8z1_70e50238ee6b479ebd61270840b4806e.webp)\\
\\
\\
flux-2-klein-4b-base-trainer\\
\\
Fine-tune FLUX.2 \[klein\] 4B from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.\\
\\
deprecated\\
\\
training](/content/models/fal-ai/flux-2-klein-4b-base-trainer/index.html)

[![Fine-tune FLUX.2 [klein] 9B from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific editing tasks.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a8b082b%2F4dsf0LE8NoXuk9Pz0Ziue_d7c1c380c4d04e03b820d06500a5749f.jpg/tr:w-1920,q-80/4dsf0LE8NoXuk9Pz0Ziue_d7c1c380c4d04e03b820d06500a5749f.webp)\\
\\
\\
flux-2-klein-9b-base-trainer\\
\\
Fine-tune FLUX.2 \[klein\] 9B from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific editing tasks.\\
\\
training](/content/models/fal-ai/flux-2-klein-9b-base-trainer/index.html)

[![Fine-tune FLUX.2 [dev] from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2Ftiger%2FnYv87OHdt503yjlNUk1P3_2551388f5f4e4537b67e8ed436333bca.jpg/tr:w-1920,q-80/nYv87OHdt503yjlNUk1P3_2551388f5f4e4537b67e8ed436333bca.webp)\\
\\
\\
flux-2-trainer-v2\\
\\
Fine-tune FLUX.2 \[dev\] from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.\\
\\
training](/content/models/fal-ai/flux-2-trainer-v2/index.html)

\\
\\
\\
z-image-trainer\\
\\
Train LoRAs on Z-Image Turbo, a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.\\
\\
turbo\\
\\
z-image\\
\\
fast\\
\\
training](/content/models/fal-ai/z-image-trainer/index.html)

\\
\\
\\
qwen-image-edit-2511-trainer\\
\\
LoRA trainer for Qwen Image Edit 2511\\
\\
deprecated\\
\\
training](/content/models/fal-ai/qwen-image-edit-2511-trainer/index.html)

[**Virtual Try On APIs**](/content/explore/virtual-try-on-apis/index.html)

Virtually try on different outfits and character styles with our collection of APIs.

\\
\\
decart/lucy2-vton/realtime\\
\\
Realtime Try On experience with Decart Lucy 2.1 VTON\\
\\
video-to-video](/content/models/decart/lucy2-vton/realtime/index.html)

\\
\\
image-apps-v2/virtual-try-on\\
\\
Try on clothes virtually by combining person and clothing images.\\
\\
fashion\\
\\
try-on\\
\\
virtual-try-on\\
\\
image-to-image](/content/models/fal-ai/image-apps-v2/virtual-try-on/index.html)

\\
\\
\\
kling/v1-5/kolors-virtual-try-on\\
\\
Kling Kolors Virtual TryOn v1.5 is a high quality image based Try-On endpoint which can be used for commercial try on.\\
\\
try-on\\
\\
fashion\\
\\
clothing\\
\\
image-to-image](/content/models/fal-ai/kling/v1-5/kolors-virtual-try-on/index.html)

\\
\\
fashn/tryon/v1.6\\
\\
FASHN v1.6 delivers precise virtual try-on capabilities, accurately rendering garment details like text and patterns at 864x1296 resolution from both on-model and flat-lay photo references.\\
\\
try-on\\
\\
fashion\\
\\
clothing\\
\\
image-to-image](/content/models/fal-ai/fashn/tryon/v1.6)

\\
\\
fashn/tryon/v1.5\\
\\
FASHN v1.5 delivers precise virtual try-on capabilities, accurately rendering garment details like text and patterns at 576x864 resolution from both on-model and flat-lay photo references.\\
\\
try-on\\
\\
fashion\\
\\
clothing\\
\\
image-to-image](/content/models/fal-ai/fashn/tryon/v1.5)

\\
\\
\\
flux-2-lora-gallery/virtual-tryon\\
\\
Virtual clothing try-on (2 images: person + garment)\\
\\
stylized\\
\\
transform\\
\\
image-to-image](/content/models/fal-ai/flux-2-lora-gallery/virtual-tryon/index.html)

\\
\\
leffa/virtual-tryon\\
\\
Leffa Virtual TryOn is a high quality image based Try-On endpoint which can be used for commercial try on.\\
\\
try-on\\
\\
fashion\\
\\
clothing\\
\\
image-to-image](/content/models/fal-ai/leffa/virtual-tryon/index.html)

\\
\\
cat-vton\\
\\
Image based high quality Virtual Try-On\\
\\
try-on\\
\\
fashion\\
\\
clothing\\
\\
image-to-image](/content/models/fal-ai/cat-vton/index.html)

[**Image to 3D Model APIs**](/content/explore/image-to-3d-model-apis/index.html)

Run the best image-to-3D models on fal

\\
\\
hunyuan-3d/v3.1/pro/image-to-3d\\
\\
Generate 3D models from images with Hunyuan 3D Pro\\
\\
3d\\
\\
hunyuan\\
\\
image-to-3d](/content/models/fal-ai/hunyuan-3d/v3.1/pro/image-to-3d/index.html)

\\
\\
trellis-2\\
\\
Generate 3D models from your images using Trellis 2. A native 3D generative model enabling versatile and high-quality 3D asset creation.\\
\\
image-to-3d\\
\\
image-to-3d](/content/models/fal-ai/trellis-2/index.html)

\\
\\
trellis\\
\\
Generate 3D models from your images using Trellis. A native 3D generative model enabling versatile and high-quality 3D asset creation.\\
\\
stylized\\
\\
image-to-3d](/content/models/fal-ai/trellis/index.html)

\\
\\
tripo3d/h3.1/image-to-3d\\
\\
Generate high-quality 3D models from a single image using Tripo H3.1.\\
\\
3d\\
\\
3d-generation\\
\\
tripo\\
\\
image-to-3d](/content/models/tripo3d/h3.1/image-to-3d/index.html)

\\
\\
meshy/v6/image-to-3d\\
\\
Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.\\
\\
image-to-3d](/content/models/fal-ai/meshy/v6/image-to-3d/index.html)

\\
\\
hyper3d/rodin/v2.5\\
\\
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. \\
\\
image-to-3d](/content/models/fal-ai/hyper3d/rodin/v2.5)

\\
\\
sam-3/3d-objects\\
\\
SAM 3D enables precise 3D reconstruction of objects from real images, while accurately reconstructing their geometry and texture.\\
\\
3d\\
\\
object\\
\\
image-to-3d](/content/models/fal-ai/sam-3/3d-objects/index.html)

\\
\\
hunyuan3d-v3/image-to-3d\\
\\
Transform your photos into ultra-high-resolution 3D models in seconds. Film-quality geometry with PBR textures, ready for games, e-commerce, and 3D printing.\\
\\
image-to-3d](/content/models/fal-ai/hunyuan3d-v3/image-to-3d/index.html)

\\
\\
tripo3d/tripo/v2.5/image-to-3d\\
\\
State of the art Image to 3D Object generation. Generate 3D model from a single image!\\
\\
stylized\\
\\
image-to-3d](/content/models/tripo3d/tripo/v2.5/image-to-3d/index.html)

\\
\\
hyper3d/rodin\\
\\
Rodin by Hyper3D generates realistic and production ready 3D models from text or images.\\
\\
stylized\\
\\
image-to-3d](/content/models/fal-ai/hyper3d/rodin/index.html)

\\
\\
tripo3d/h3.1/multiview-to-3d\\
\\
Generate 3D models from multiple view images using Tripo H3.1.\\
\\
3d\\
\\
multiview-to-3d\\
\\
3d-generation\\
\\
image-to-3d](/content/models/tripo3d/h3.1/multiview-to-3d/index.html)

\\
\\
meshy/v6/multi-image-to-3d\\
\\
Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.\\
\\
image-to-3d](/content/models/fal-ai/meshy/v6/multi-image-to-3d/index.html)

\\
\\
hyper3d/rodin/v2\\
\\
Rodin by Hyper3D generates realistic and production ready 3D models from text or images.\\
\\
text-to-3d\\
\\
image-to-3d](/content/models/fal-ai/hyper3d/rodin/v2/index.html)

\\
\\
hunyuan3d/v2\\
\\
Generate 3D models from your images using Hunyuan 3D. A native 3D generative model enabling versatile and high-quality 3D asset creation.\\
\\
stylized\\
\\
image-to-3d](/content/models/fal-ai/hunyuan3d/v2/index.html)

\\
\\
sam-3/3d-body\\
\\
SAM 3D allows for accurate 3D reconstruction of human body shape and position from a single image.\\
\\
3d\\
\\
human\\
\\
pose\\
\\
image-to-3d](/content/models/fal-ai/sam-3/3d-body/index.html)

\\
\\
tripo3d/triposplat\\
\\
TripoSplat is an open-source model from TripoAI / VAST AI Research that converts a single 2D image into high-quality 3D Gaussians using a novel learned density-control approach\\
\\
3d\\
\\
gaussian-splat\\
\\
image-to-3d](/content/models/tripo3d/triposplat/index.html)

\\
\\
hunyuan-3d/v3.1/rapid/image-to-3d\\
\\
Rapidly generate 3D models from images using Hunyuan 3D.\\
\\
3d\\
\\
hunyuan\\
\\
image-to-3d](/content/models/fal-ai/hunyuan-3d/v3.1/rapid/image-to-3d/index.html)

\\
\\
meshy/v6-preview/image-to-3d\\
\\
Meshy-6-Preview is the latest model from Meshy. It generates realistic and production ready 3D models.\\
\\
image-to-3d](/content/models/fal-ai/meshy/v6-preview/image-to-3d/index.html)

\\
\\
trellis/multi\\
\\
Generate 3D models from multiple images using Trellis. A native 3D generative model enabling versatile and high-quality 3D asset creation.\\
\\
stylized\\
\\
image-to-3d](/content/models/fal-ai/trellis/multi/index.html)

\\
\\
hyper3d/rodin/v2.5/fast\\
\\
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model. \\
\\
image-to-3d](/content/models/fal-ai/hyper3d/rodin/v2.5/fast/index.html)

\\
\\
hunyuan3d/v2/multi-view\\
\\
Generate 3D models from your images using Hunyuan 3D. A native 3D generative model enabling versatile and high-quality 3D asset creation.\\
\\
stylized\\
\\
image-to-3d](/content/models/fal-ai/hunyuan3d/v2/multi-view/index.html)

\\
\\
pixal3d\\
\\
Pixal3D turns a single image into a high-fidelity 3D model with detailed geometry and realistic textures.\\
\\
stylized\\
\\
transform\\
\\
image-to-3d](/content/models/fal-ai/pixal3d/index.html)

\\
\\
tripo3d/p1/image-to-3d\\
\\
Generate 3D models from a single image using Tripo P1.\\
\\
3d\\
\\
3d-generation\\
\\
tripo\\
\\
image-to-3d](/content/models/tripo3d/p1/image-to-3d/index.html)

\\
\\
tripo3d/tripo/v2.5/multiview-to-3d\\
\\
State of the art Multiview to 3D Object generation. Generate 3D models from multiple images!\\
\\
stylized\\
\\
multiview\\
\\
image-to-3d](/content/models/tripo3d/tripo/v2.5/multiview-to-3d/index.html)

\\
\\
triposr\\
\\
State of the art Image to 3D Object generation\\
\\
image-to-3d](/content/models/fal-ai/triposr/index.html)

\\
\\
hunyuan3d/v2/turbo\\
\\
Generate 3D models from your images using Hunyuan 3D. A native 3D generative model enabling versatile and high-quality 3D asset creation.\\
\\
stylized\\
\\
image-to-3d](/content/models/fal-ai/hunyuan3d/v2/turbo/index.html)

\\
\\
hunyuan\_world/image-to-world\\
\\
Hunyuan World 1.0 turns a single image into a panorama or a 3D world. It creates realistic scenes from the image, allowing you to explore and view it from different angles.\\
\\
image-to-3d](/content/models/fal-ai/hunyuan_world/image-to-world/index.html)

\\
\\
reconviagen-0.5\\
\\
Generate 3D models from one or more images using ReconViaGen 0.5\\
\\
multi-view\\
\\
3d-reconstruction\\
\\
image-to-3d](/content/models/fal-ai/reconviagen-0.5)

[**Text to 3D Model APIs**](/content/explore/text-to-3d-model-apis/index.html)

This is our collection of the best text-to-3D model APIs available on fal.

\\
\\
meshy/v6/text-to-3d\\
\\
Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.\\
\\
text-to-3d](/content/models/fal-ai/meshy/v6/text-to-3d/index.html)

\\
\\
hunyuan-3d/v3.1/pro/text-to-3d\\
\\
Generate 3D models from text prompts with Hunyuan 3D Pro\\
\\
3d\\
\\
hunyuan\\
\\
text-to-3d](/content/models/fal-ai/hunyuan-3d/v3.1/pro/text-to-3d/index.html)

\\
\\
tripo3d/h3.1/text-to-3d\\
\\
Generate 3D models from text descriptions using Tripo H3.1.\\
\\
3d\\
\\
3d-generation\\
\\
tripo\\
\\
text-to-3d](/content/models/tripo3d/h3.1/text-to-3d/index.html)

\\
\\
hunyuan-motion\\
\\
Generate 3D human motions via text-to-generation interface of Hunyuan Motion!\\
\\
motion\\
\\
text-to-3d](/content/models/fal-ai/hunyuan-motion/index.html)

\\
\\
hunyuan-3d/v3.1/rapid/text-to-3d\\
\\
Create detailed, fully-textured 3D models with text\\
\\
3d\\
\\
text-to-3d](/content/models/fal-ai/hunyuan-3d/v3.1/rapid/text-to-3d/index.html)

\\
\\
hunyuan3d-v3/text-to-3d\\
\\
Turn simple sketches into detailed, fully-textured 3D models. Instantly convert your concept designs into formats ready for Unity, Unreal, and Blender.\\
\\
text-to-3d](/content/models/fal-ai/hunyuan3d-v3/text-to-3d/index.html)

\\
\\
hyper3d/rodin/v2.5/text-to-3d\\
\\
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. \\
\\
text-to-3d](/content/models/fal-ai/hyper3d/rodin/v2.5/text-to-3d/index.html)

\\
\\
meshy/v6-preview/text-to-3d\\
\\
Meshy-6-Preview is the latest model from Meshy. It generates realistic and production ready 3D models.\\
\\
text-to-3d](/content/models/fal-ai/meshy/v6-preview/text-to-3d/index.html)

\\
\\
hyper3d/rodin/v2.5/text-to-3d/fast\\
\\
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model. \\
\\
text-to-3d](/content/models/fal-ai/hyper3d/rodin/v2.5/text-to-3d/fast/index.html)

\\
\\
tripo3d/p1/text-to-3d\\
\\
Generate 3D models from text descriptions using Tripo P1.\\
\\
3d\\
\\
3d-generation\\
\\
tripo\\
\\
text-to-3d](/content/models/tripo3d/p1/text-to-3d/index.html)

\\
\\
hunyuan-motion/fast\\
\\
Generate 3D human motions via text-to-generation interface of Hunyuan Motion!\\
\\
motion\\
\\
text-to-3d](/content/models/fal-ai/hunyuan-motion/fast/index.html)

[**Best Utility Models**](/content/explore/best-utility-models/index.html)

Specialized models for supporting tasks like background removal, nsfw detection, upscaling and much more.

\\
\\
x-ailab/nsfw\\
\\
Predict whether an image is NSFW or SFW.\\
\\
filter\\
\\
safety\\
\\
utility\\
\\
vision](/content/models/fal-ai/x-ailab/nsfw/index.html)

\\
\\
topaz/upscale/image\\
\\
Use the powerful and accurate topaz image enhancer to enhance your images.\\
\\
image-to-image](/content/models/fal-ai/topaz/upscale/image/index.html)

\\
\\
\\
bria/video/background-removal\\
\\
Automatically remove backgrounds from videos -perfect for creating clean, professional content without a green screen.\\
\\
background-removal\\
\\
video-to-video](/content/models/bria/video/background-removal/index.html)

\\
\\
topaz/upscale/video\\
\\
Professional-grade video upscaling using Topaz technology. Enhance your videos with high-quality upscaling.\\
\\
upscaling\\
\\
high-res\\
\\
video-to-video](/content/models/fal-ai/topaz/upscale/video/index.html)

\\
\\
\\
bria/video/background-removal/realtime\\
\\
Remove video backgrounds in real time with Bria’s VRMBG 3.0 model. Built for live streaming, real-time video apps, content creation, and low-latency workflows that need fast, accurate background removal.\\
\\
bria\\
\\
video\\
\\
background-removal\\
\\
video-to-video](/content/models/bria/video/background-removal/realtime/index.html)

[**Text To Image APIs**](/content/explore/text-to-image-apis/index.html)

Use the latest state of the art text to image model APIs

\\
\\
fast-sdxl\\
\\
Run SDXL at the speed of light\\
\\
diffusion\\
\\
lora\\
\\
embeddings\\
\\
text-to-image](/content/models/fal-ai/fast-sdxl/index.html)
