# Real-Time Inference

Real-time inference uses WebSockets for persistent connections, enabling sub-100ms image generation. This is ideal for interactive applications like real-time creativity tools and camera-based inputs.

Unlike [queue-based inference](/docs/documentation/model-apis/inference), real-time connections bypass the queue entirely and route inputs directly to a runner. This eliminates queue wait time, and because the WebSocket maintains a persistent connection, the runner stays warm for all subsequent messages after the initial connection. The first connection may still incur a cold start if no runner is already available. Only models with an explicit real-time endpoint are supported.

> **Warning**  
> Only models that explicitly support real-time inference can be used with the realtime client. Standard queue-based models do not have a realtime endpoint.

> **Note**  
> If the model you want has no realtime endpoint but you still want a persistent connection, see [HTTP over WebSockets](/docs/documentation/model-apis/inference/websockets) — it carries a model's ordinary HTTP request and response format over `wss://ws.fal.run/{model_id}`, and works with any endpoint.

## Supported Models

| Model Title                  | Description                              |
|------------------------------|------------------------------------------|
| [fast-lcm-diffusion](/content/models/fal-ai/fast-lcm-diffusion/index.html)    | SDXL with Latent Consistency Models     |
| [fast-turbo-diffusion](/content/models/fal-ai/fast-turbo-diffusion/index.html) | Optimized SDXL Turbo                     |

## Quick Start

### JavaScript Example
```javascript
import { fal } from "@fal-ai/client";

const connection = fal.realtime.connect("fal-ai/fast-lcm-diffusion", {
  onResult: (result) => {
    console.log(result);
  },
  onError: (error) => {
    console.error(error);
  },
});

connection.send({
  prompt: "a sunset over mountains",
  sync_mode: true,
  image_url: "data:image/png;base64,..."
});
```

### Python Example
```python
import fal_client

with fal_client.realtime("fal-ai/fast-lcm-diffusion") as connection:
    connection.send({
        "prompt": "a sunset over mountains",
        "sync_mode": True,
        "image_url": "data:image/png;base64,..."
    })
    result = connection.recv()
    print(result)
```

### Async Python Example
```python
import asyncio
import fal_client

async def realtime():
    async with fal_client.realtime_async("fal-ai/fast-lcm-diffusion") as connection:
        await connection.send({
            "prompt": "a sunset over mountains",
            "sync_mode": True,
            "image_url": "data:image/png;base64,..."
        })
        result = await connection.recv()
        print(result)

asyncio.run(realtime())
```

## Performance Tips

For the fastest inference:

* Use **512x512** input dimensions (fastest)
* Provide images as base64 encoded data URLs
* Set `sync_mode: true` to receive base64 encoded responses
* 768x768 and 1024x1024 also work well, but 512x512 is optimal

## Keeping API Keys Secure

WebSocket connections from browsers cannot safely embed API keys. There are two approaches for client-side authentication: a proxy URL or a token provider.

### Proxy URL

The simplest approach. Point the client at a server-side proxy that adds your API key:

```javascript
import { fal } from "@fal-ai/client";

fal.config({
  proxyUrl: "/api/fal/proxy",
});

const connection = fal.realtime.connect("fal-ai/fast-lcm-diffusion", {
  connectionKey: "realtime-demo",
  throttleInterval: 128,
  onResult(result) {
    // handle result
  },
});
```

### Token Provider

For more control, use a `tokenProvider` function that fetches short-lived JWT tokens from your backend. This is useful when you need per-user authentication or want to restrict which apps a token can access.

> **Warning**  
> **Protect your token endpoint with authentication.** The endpoint that generates fal tokens should verify that the request comes from an authenticated user in your application. Without proper authentication, anyone could use your endpoint to generate tokens and consume your fal credits.

**Client-side example:**
```typescript
import { fal, type TokenProvider } from "@fal-ai/client";

const myTokenProvider: TokenProvider = async (app) => {
  const response = await fetch(`/api/fal/token?app=${encodeURIComponent(app)}`);
  const { token } = await response.json();
  return token;
};

const connection = fal.realtime.connect("fal-ai/fast-lcm-diffusion", {
  tokenProvider: myTokenProvider,
  tokenExpirationSeconds: 120,
  onResult: (result) => {
    console.log(result);
  },
});

connection.send({
  prompt: "a cat",
  sync_mode: true,
});
```

## Differences from Queue-Based Inference

| Parameter         | Behavior with Real-Time                                        |
| ----------------- | -------------------------------------------------------------- |
| `start_timeout`   | No effect. There is no queue wait                              |
| `priority`        | No effect. No queue ordering                                   |
| `webhook_url`     | Not supported. Results stream back over the WebSocket          |
| Automatic retries | Not available. Failed messages return errors on the connection |
| `X-Fal-No-Retry`  | No effect. No retry mechanism to disable                       |

## Custom WebSocket Path

By default, the realtime client connects to the `/realtime` path on the app (e.g., `wss://fal.run/fal-ai/my-app/realtime`). If your app exposes a realtime endpoint at a different path, use the `path` option:

```typescript
const connection = fal.realtime.connect("fal-ai/my-app", {
  path: "/my-custom-ws",
  onResult: (result) => console.log(result),
});
```

## Realtime vs Streaming

| Feature        | Realtime (WebSocket)                    | Streaming (SSE)                   |
| -------------- | --------------------------------------- | --------------------------------- |
| **Direction**  | Bidirectional (client and server)       | One-way (server to client)        |
| **Connection** | Persistent, reusable                    | New connection per request        |
| **Latency**    | Lower (connection reuse)                | Higher (new connection each time) |
| **Best for**   | Interactive apps, back-to-back requests | Progressive output, previews      |
| **Protocol**   | Binary msgpack (default, customizable)  | JSON over SSE                     |

## Protocol Details

The realtime client uses [msgpack](https://msgpack.org/) for binary serialization by default across all SDKs, which is more efficient than JSON for transmitting image data. In Python, `realtime()` and `realtime_async()` provide a `RealtimeConnection` with `send()` and `recv()` methods. In JavaScript, `fal.realtime.connect()` uses callback-based `onResult` and `onError` handlers.

## Video Tutorial

Build a Real-Time AI Image App with WebSockets, Next.js, and fal.ai:
