# nvidia/cosmos-3-super/image-to-video

## Image to Video

Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs.

### Inference

### Commercial use

### About

Generate a video from a first-frame image and prompt using Cosmos3-Super via vLLM-Omni.

### 1. Calling the API

#### Install the client

The client provides a convenient way to interact with the model API.

```bash
npm install --save @fal-ai/client
```

##### Migrate to @fal-ai/client

The `@fal-ai/serverless-client` package has been deprecated in favor of `@fal-ai/client`. Install the new package and update your imports — see [client setup](/content/docs/documentation/model-apis/inference/client-setup#installation/index.html).

#### Setup your API Key

Set `FAL_KEY` as an environment variable in your runtime.

```bash
export FAL_KEY="YOUR_API_KEY"
```

#### Submit a request

The client API handles the API submit protocol. It will handle the request status updates and return the result when the request is completed.

```javascript
import { fal } from "@fal-ai/client";

const result = await fal.subscribe("nvidia/cosmos-3-super/image-to-video", {
  input: {
    prompt: "The camera slowly pushes in as the subject turns their head toward the light, hair drifting in a gentle breeze, dust motes floating through warm afternoon sun.",
    image_url: "https://storage.googleapis.com/falserverless/example_inputs/hunyuan_i2v.jpg"
  },
  logs: true,
  onQueueUpdate: (update) => {
    if (update.status === "IN_PROGRESS") {
      update.logs.map((log) => log.message).forEach(console.log);
    }
  },
});
console.log(result.data);
console.log(result.requestId);
```

## 2. Authentication

The API uses an API Key for authentication. It is recommended you set the `FAL_KEY` environment variable in your runtime when possible.

### API Key

In case your app is running in an environment where you cannot set environment variables, you can set the API Key manually as a client configuration.

```javascript
import { fal } from "@fal-ai/client";

fal.config({
  credentials: "YOUR_FAL_KEY"
});
```

##### Protect your API Key

When running code on the client-side (e.g. in a browser, mobile app or GUI applications), make sure to not expose your `FAL_KEY`. Instead, **use a server-side proxy** to make requests to the API. For more information, check out our [server-side integration guide](/content/docs/documentation/model-apis/inference/proxy-setup/index.html).

## 3. Queue

##### Long-running requests

For long-running requests, such as _training_ jobs or models with slower inference times, it is recommended to check the [Queue](/content/docs/documentation/model-apis/inference/queue/index.html) status and rely on [Webhooks](/content/docs/documentation/model-apis/inference/webhooks/index.html) instead of blocking while waiting for the result.

### Submit a request

The client API provides a convenient way to submit requests to the model.

```javascript
import { fal } from "@fal-ai/client";

const { request_id } = await fal.queue.submit("nvidia/cosmos-3-super/image-to-video", {
  input: {
    prompt: "The camera slowly pushes in as the subject turns their head toward the light, hair drifting in a gentle breeze, dust motes floating through warm afternoon sun.",
    image_url: "https://storage.googleapis.com/falserverless/example_inputs/hunyuan_i2v.jpg"
  },
  webhookUrl: "https://optional.webhook.url/for/results",
});
```

### Fetch request status

You can fetch the status of a request to check if it is completed or still in progress.

```javascript
import { fal } from "@fal-ai/client";

const status = await fal.queue.status("nvidia/cosmos-3-super/image-to-video", {
  requestId: "764cabcf-b745-4b3e-ae38-1200304cf45b",
  logs: true,
});
```

### Get the result

Once the request is completed, you can fetch the result. See the [Output Schema](/content/models/nvidia/cosmos-3-super/image-to-video/api#schema-output/index.html) for the expected result format.

```javascript
import { fal } from "@fal-ai/client";

const result = await fal.queue.result("nvidia/cosmos-3-super/image-to-video", {
  requestId: "764cabcf-b745-4b3e-ae38-1200304cf45b"
});
console.log(result.data);
console.log(result.requestId);
```

## 4. Files

Some attributes in the API accept file URLs as input. Whenever that's the case you can pass your own URL or a Base64 data URI.

### Data URI (base64)

You can pass a Base64 data URI as a file input. The API will handle the file decoding for you. Keep in mind that for large files, this alternative although convenient can impact the request performance.

### Hosted files (URL)

You can also pass your own URLs as long as they are publicly accessible. Be aware that some hosts might block cross-site requests, rate-limit, or consider the request as a bot.

### Uploading files

We provide a convenient file storage that allows you to upload files and use them in your requests. You can upload files using the client API and use the returned URL in your requests.

```javascript
import { fal } from "@fal-ai/client";

const file = new File(["Hello, World!"], "hello.txt", { type: "text/plain" });
const url = await fal.storage.upload(file);
```

##### Auto uploads

The client will auto-upload the file for you if you pass a binary object (e.g. `File`, `Data`).

Read more about file handling in our [file upload guide](/content/docs/documentation/model-apis/fal-cdn#uploading-files/index.html).

## 5. Schema

### Input

- `prompt` `string`\* required: Text prompt describing the motion and scene of the video to generate.
- `image_url` `string`\* required: URL of the conditioning first-frame image for the video.
- `negative_prompt` `string`: Content to steer the generation away from (artifacts, unwanted motion). Defaults to NVIDIA's recommended i2v negative prompt; pass an empty string to disable.
- `enable_prompt_expansion` `boolean`: If true, the Cosmos3-Nano Reasoner rewrites the prompt into the dense caption Cosmos3 was trained on.
- `agentic_max_iterations` `integer`: Maximum agentic prompt stages when agentic generation is enabled.
- `agentic_samples_per_iteration` `integer`: Candidate videos to generate and judge per agentic iteration.
- `agentic_early_stop` `boolean`: Stop the agentic loop early when the critic score clears the strict quality threshold.
- `image_size` `ImageSize | Enum`: The size of the generated video.
- `num_frames` `integer`: Number of frames to generate.
- `frames_per_second` `integer`: Frames per second of the output video.
- `num_inference_steps` `integer`: Number of denoising steps.
- `guidance_scale` `float`: Classifier-free guidance scale.
- `seed` `integer`: The same seed and prompt given to the same model version will produce the same video every time.
- `enable_safety_checker` `boolean`: Enable content moderation for the input prompt and image.
- `sync_mode` `boolean`: If `True`, the video is returned as a data URI and the output data won't be available in the request history.

### Output

- `video` `VideoFile`\* required: The generated video.
- `seed` `integer`\* required: The seed used for generation.

### Other types
#### ImageFile
- `url` `string`\* required: The URL where the file can be downloaded from.
- `content_type` `string`: The mime type of the file.
- `file_name` `string`: The name of the file.
- `file_size` `integer`: The size of the file in bytes.
- `width` `integer`: The width of the image
- `height` `integer`: The height of the image

#### ImageSize
- `width` `integer`: The width of the generated image.
- `height` `integer`: The height of the generated image.

#### VideoFile
- `url` `string`\* required: The URL where the file can be downloaded from.
- `content_type` `string`: The mime type of the file.
- `file_name` `string`: The name of the file.
- `file_size` `integer`: The size of the file in bytes.
- `width` `integer`: The width of the video.
- `height` `integer`: The height of the video.
- `fps` `float`: The FPS of the video.
- `duration` `float`: The duration of the video.
- `num_frames` `integer`: The number of frames in the video.
