atlascloud/infinitetalk

Ses-Video

InfiniteTalk API by Atlas Cloud

atlascloud/infinitetalk

Infinitetalk

InfiniteTalk turns a reference portrait and audio into a realistic talking-head video with lip-sync, supporting up to 10-minute audio in 480p or 720p.

Girdi

Prompt

Görüntü *

Dosyaları buraya sürükleyip bırakabilir veya yüklemek için tıklayabilirsiniz

MAX:1

Ses *

Dosyaları buraya sürükleyip bırakabilir veya yüklemek için tıklayabilirsiniz

MAX:1

Çözünürlük

Seed

Çıktı

Boşta

Oluşturulan videolarınız burada görünecek

Parametreleri yapılandırın ve oluşturmaya başlamak için Çalıştır'a tıklayın

Her çalıştırma $0.03 maliyete sahip. 10$ ile yaklaşık 333 kez çalıştırabilirsiniz.

Şununla devam edebilirsiniz:

Seedance 2.0 Kling v3 Vidu Wan2.7

Parametreler

Kod örneği
import requests
import time

# Step 1: Start video generation
generate_url = "https://api.atlascloud.ai/api/v1/model/generateVideo"
headers = {
    "Content-Type": "application/json",
    "Authorization": "Bearer $ATLASCLOUD_API_KEY"
}
data = {
    "model": "atlascloud/infinitetalk",  # Required. Model name
    "image": "example_value",  # Required. Portrait photo URL
    "audio": "example_value",  # Required. Driving audio file URL (WAV or MP3)
    "prompt": "A beautiful sunset over the ocean with gentle waves",  # Text guidance for expression, posture, or behavior
    "resolution": "480p",  # Output resolution. options: 480p | 720p
    "seed": -1,  # Random seed for reproducibility
}

generate_response = requests.post(generate_url, headers=headers, json=data)
generate_result = generate_response.json()
prediction_id = generate_result["data"]["id"]

# Step 2: Poll for result
poll_url = f"https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}"

def check_status():
    while True:
        response = requests.get(poll_url, headers={"Authorization": "Bearer $ATLASCLOUD_API_KEY"})
        result = response.json()

        if result["data"]["status"] in ["completed", "succeeded"]:
            print("Generated video:", result["data"]["outputs"][0])
            return result["data"]["outputs"][0]
        elif result["data"]["status"] == "failed":
            raise Exception(result["data"]["error"] or "Generation failed")
        else:
            # Still processing, wait 2 seconds
            time.sleep(2)

video_url = check_status()

Kurulum

Programlama diliniz için gerekli paketi kurun.

pip install requests

Kimlik Doğrulama

Tüm API istekleri, API anahtarı ile kimlik doğrulama gerektirir. API anahtarınızı Atlas Cloud kontrol panelinden alabilirsiniz.

export ATLASCLOUD_API_KEY="your-api-key-here"

HTTP Başlıkları

import os

API_KEY = os.environ.get("ATLASCLOUD_API_KEY")
headers = {
    "Content-Type": "application/json",
    "Authorization": f"Bearer {API_KEY}"
}

API anahtarınızı güvende tutun

API anahtarınızı asla istemci tarafı kodunda veya herkese açık depolarda ifşa etmeyin. Bunun yerine ortam değişkenleri veya arka uç proxy kullanın.

İstek gönder

import requests

url = "https://api.atlascloud.ai/api/v1/model/generateVideo"
headers = {
    "Content-Type": "application/json",
    "Authorization": "Bearer $ATLASCLOUD_API_KEY"
}
data = {
    "model": "your-model",
    "prompt": "A beautiful landscape"
}

response = requests.post(url, headers=headers, json=data)
print(response.json())

İstek Gönder

Asenkron bir oluşturma isteği gönderin. API, durumu kontrol etmek ve sonucu almak için kullanabileceğiniz bir tahmin ID'si döndürür.

POST/api/v1/model/generateVideo

İstek Gövdesi

import requests

url = "https://api.atlascloud.ai/api/v1/model/generateVideo"
headers = {
    "Content-Type": "application/json",
    "Authorization": "Bearer $ATLASCLOUD_API_KEY"
}

data = {
    "model": "atlascloud/infinitetalk",
    "prompt": "A beautiful sunset over the ocean with gentle waves"
}

response = requests.post(url, headers=headers, json=data)
result = response.json()

print(f"Prediction ID: {result['data']['id']}")
print(f"Status: {result['data']['status']}")

Yanıt

{
  "code": 200,
  "data": {
    "id": "pred_abc123",
    "status": "processing",
    "model": "model-name",
    "created_at": "2025-01-01T00:00:00Z"
  }
}

Durumu Kontrol Et

İsteğinizin mevcut durumunu kontrol etmek için tahmin uç noktasını sorgulayın.

GET/api/v1/model/prediction/{prediction_id}

Sorgulama Örneği

import requests
import time

prediction_id = "pred_abc123"
url = f"https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}"
headers = { "Authorization": "Bearer $ATLASCLOUD_API_KEY" }

while True:
    response = requests.get(url, headers=headers)
    result = response.json()
    status = result["data"]["status"]
    print(f"Status: {status}")

    if status in ["completed", "succeeded"]:
        output_url = result["data"]["outputs"][0]
        print(f"Output URL: {output_url}")
        break
    elif status == "failed":
        print(f"Error: {result['data'].get('error', 'Unknown')}")
        break

    time.sleep(3)

Durum Değerleri

processingİstek hâlâ işleniyor.

completedOluşturma tamamlandı. Çıktılar kullanılabilir.

succeededOluşturma başarılı oldu. Çıktılar kullanılabilir.

failedOluşturma başarısız oldu. Hata alanını kontrol edin.

Tamamlanmış Yanıt

{
  "data": {
    "id": "pred_abc123",
    "status": "completed",
    "outputs": [
      "https://storage.atlascloud.ai/outputs/result.mp4"
    ],
    "metrics": {
      "predict_time": 45.2
    },
    "created_at": "2025-01-01T00:00:00Z",
    "completed_at": "2025-01-01T00:00:10Z"
  }
}

Dosya Yükle

Dosyaları Atlas Cloud depolama alanına yükleyin ve API isteklerinizde kullanabileceğiniz bir URL alın. Yüklemek için multipart/form-data kullanın.

POST/api/v1/model/uploadMedia

Yükleme Örneği

import requests

url = "https://api.atlascloud.ai/api/v1/model/uploadMedia"
headers = { "Authorization": "Bearer $ATLASCLOUD_API_KEY" }

with open("image.png", "rb") as f:
    files = {"file": ("image.png", f, "image/png")}
    response = requests.post(url, headers=headers, files=files)

result = response.json()
download_url = result["data"]["download_url"]
print(f"File URL: {download_url}")

Yanıt

{
  "data": {
    "download_url": "https://storage.atlascloud.ai/uploads/abc123/image.png",
    "file_name": "image.png",
    "content_type": "image/png",
    "size": 1024000
  }
}

Input Schema

İstek gövdesinde aşağıdaki parametreler kabul edilir.

Toplam: 6Zorunlu: 3İsteğe Bağlı: 3

modelstringrequired

Model name.

Default: "atlascloud/infinitetalk"

imagestringrequired

Portrait photo URL.

audiostringrequired

Driving audio file URL (WAV or MP3).

promptstring

Text guidance for expression, posture, or behavior.

resolutionstring

Output resolution.

Default: "480p"

480p720p

seedinteger

Random seed for reproducibility. -1 for random.

Default: -1

Örnek İstek Gövdesi

{
  "model": "atlascloud/infinitetalk",
  "image": "example_image",
  "audio": "example_audio",
  "resolution": "480p",
  "seed": -1
}

Output Schema

API, oluşturulan çıktı URL'lerini içeren bir tahmin yanıtı döndürür.

idstringrequired

Unique identifier for the prediction.

statusstringrequired

Current status of the prediction.

processingcompletedsucceededfailed

modelstringrequired

The model used for generation.

outputsarray[string]

Array of output URLs. Available when status is "completed".

errorstring

Error message if status is "failed".

metricsobject

Performance metrics.

predict_timenumber

Time taken for video generation in seconds.

created_atstringrequired

ISO 8601 timestamp when the prediction was created.

Format: date-time

completed_atstring

ISO 8601 timestamp when the prediction was completed.

Format: date-time

Örnek Yanıt

{
  "id": "pred_abc123",
  "status": "completed",
  "model": "model-name",
  "outputs": [
    "https://storage.atlascloud.ai/outputs/result.mp4"
  ],
  "metrics": {
    "predict_time": 45.2
  },
  "created_at": "2025-01-01T00:00:00Z",
  "completed_at": "2025-01-01T00:00:10Z"
}

Atlas Cloud Skills

Atlas Cloud Skills, 400'den fazla AI modelini doğrudan AI kodlama asistanınıza entegre eder. Kurmak için tek bir komut, ardından görüntü, video oluşturmak ve LLM ile sohbet etmek için doğal dil kullanın.

Desteklenen İstemciler

Claude Code

OpenAI Codex

Gemini CLI

Cursor

Windsurf

VS Code

Trae

GitHub Copilot

Cline

Roo Code

Amp

Goose

Replit

40+ desteklenen i̇stemciler

Kurulum

npx skills add AtlasCloudAI/atlas-cloud-skills

API Anahtarını Ayarla

API anahtarınızı Atlas Cloud kontrol panelinden alın ve ortam değişkeni olarak ayarlayın.

export ATLASCLOUD_API_KEY="your-api-key-here"

Yetenekler

Kurulduktan sonra, tüm Atlas Cloud modellerine erişmek için AI asistanınızda doğal dil kullanabilirsiniz.

Görüntü OluşturmaNano Banana 2, Z-Image ve daha fazla model ile görüntüler oluşturun.

Video OluşturmaKling, Vidu, Veo vb. ile metin veya görüntülerden videolar oluşturun.

LLM SohbetQwen, DeepSeek ve diğer büyük dil modelleri ile sohbet edin.

Medya YüklemeGörüntü düzenleme ve görüntüden videoya iş akışları için yerel dosyaları yükleyin.

Daha fazla bilgi

github.com/AtlasCloudAI/atlas-cloud-skills

MCP Server

Atlas Cloud MCP Server, IDE'nizi Model Context Protocol aracılığıyla 400'den fazla AI modeline bağlar. Herhangi bir MCP uyumlu istemci ile çalışır.

Desteklenen İstemciler

Cursor

VS Code

Windsurf

Claude Code

OpenAI Codex

Gemini CLI

Cline

Roo Code

100+ desteklenen i̇stemciler

Kurulum

npx -y atlascloud-mcp

Yapılandırma

Aşağıdaki yapılandırmayı IDE'nizin MCP ayarları dosyasına ekleyin.

{
  "mcpServers": {
    "atlascloud": {
      "command": "npx",
      "args": [
        "-y",
        "atlascloud-mcp"
      ],
      "env": {
        "ATLASCLOUD_API_KEY": "your-api-key-here"
      }
    }
  }
}

Mevcut Araçlar

atlas_generate_imageMetin istemlerinden görüntüler oluşturun.

atlas_generate_videoMetin veya görüntülerden videolar oluşturun.

atlas_chatBüyük dil modelleri ile sohbet edin.

atlas_list_models400'den fazla mevcut AI modelini keşfedin.

atlas_quick_generateOtomatik model seçimi ile tek adımda içerik oluşturma.

atlas_upload_mediaAPI iş akışları için yerel dosyaları yükleyin.

Daha fazla bilgi

github.com/AtlasCloudAI/mcp-server

API Şeması

{
  "components": {
    "schemas": {
      "Input": {
        "properties": {
          "model": {
            "type": "string",
            "description": "Model name.",
            "default": "atlascloud/infinitetalk"
          },
          "image": {
            "type": "string",
            "description": "Portrait photo URL."
          },
          "audio": {
            "type": "string",
            "description": "Driving audio file URL (WAV or MP3)."
          },
          "prompt": {
            "type": "string",
            "description": "Text guidance for expression, posture, or behavior."
          },
          "resolution": {
            "type": "string",
            "description": "Output resolution.",
            "default": "480p",
            "enum": [
              "480p",
              "720p"
            ],
            "x-ui-component": "select"
          },
          "seed": {
            "type": "integer",
            "description": "Random seed for reproducibility. -1 for random.",
            "default": -1
          }
        },
        "required": [
          "model",
          "image",
          "audio"
        ],
        "type": "object",
        "x-order-properties": [
          "model",
          "prompt",
          "image",
          "audio",
          "resolution",
          "seed"
        ]
      },
      "Output": {
        "properties": {
          "outputs": {
            "type": "array",
            "description": "List of generated video URLs.",
            "items": {
              "type": "string"
            }
          }
        },
        "type": "object"
      }
    }
  }
}

LLM Dostu İstem Şablonu

# atlascloud/infinitetalk

> InfiniteTalk turns a reference portrait and audio into a realistic talking-head video with lip-sync, supporting up to 10-minute audio in 480p or 720p.


## Overview

- **Submit endpoint (POST)**: `https://api.atlascloud.ai/api/v1/model/generateVideo` — start an async generation; returns a `prediction_id`
- **Poll endpoint (GET)**: `https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}` — poll this until the prediction finishes
- **Model ID**: `atlascloud/infinitetalk`


## API Information

This model can be used via our HTTP API or more conveniently via our client libraries.
See the input and output schema below, as well as the usage examples.


### Input Schema

The API accepts the following input parameters:

- **`model`** (`string`, _required_):
  Model name.
  - Default: `"atlascloud/infinitetalk"`

- **`prompt`** (`string`, _optional_):
  Text guidance for expression, posture, or behavior.

- **`image`** (`string`, _required_):
  Portrait photo URL.

- **`audio`** (`string`, _required_):
  Driving audio file URL (WAV or MP3).

- **`resolution`** (`string`, _optional_):
  Output resolution.
  - Default: `"480p"`
  - Options: "480p", "720p"

- **`seed`** (`integer`, _optional_):
  Random seed for reproducibility. -1 for random.
  - Default: `-1`



**Required Parameters Example**:

```json
{
  "model": "atlascloud/infinitetalk",
  "image": "",
  "audio": ""
}
```


**Full Example**:

```json
{
  "model": "atlascloud/infinitetalk",
  "prompt": "",
  "image": "",
  "audio": "",
  "resolution": "480p",
  "seed": -1
}
```


## Usage Examples

### cURL

```bash
# Step 1: Start generation (async)
curl -X POST "https://api.atlascloud.ai/api/v1/model/generateVideo" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "atlascloud/infinitetalk",
  "prompt": "",
  "image": "",
  "audio": "",
  "resolution": "480p",
  "seed": -1
}'

# Response will contain: {"code": 200, "data": {"id": "prediction_id", "status": "processing"}}

# Step 2: Poll for result (replace {prediction_id} with the id returned above)
curl -X GET "https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY"

# Keep polling until status is "completed", "succeeded" or "failed"
# When completed, outputs will contain the generated content URL(s)
```

## Additional Resources

### Documentation

- [Model Playground](https://www.atlascloud.ai/models/atlascloud/infinitetalk)

A young woman sitting at a desk, holding a Bluetooth speaker, talking to camera with a smile.

Yükleniyor...

InfiniteTalk: Audio-Driven Talking Video Generation

1. Introduction

InfiniteTalk is an audio-driven video generation model developed by AtlasCloud that transforms a single portrait image into a realistic talking-head video synchronized to any speech audio input. Built on a modified Wan2.1 I2V-14B diffusion transformer backbone with a dedicated audio cross-attention module, InfiniteTalk achieves phoneme-level lip synchronization while preserving the subject's identity, hairstyle, clothing, and background throughout the entire video.

InfiniteTalk's core innovation lies in its triple cross-attention architecture: each transformer block processes visual self-attention, text prompt cross-attention, and frame-level audio cross-attention in sequence, enabling precise per-frame audio-visual alignment. Combined with a streaming inference pipeline that processes video in overlapping segments, InfiniteTalk supports continuous video generation of up to 10 minutes from a single request — far exceeding the typical 5–15 second limit of conventional image-to-video models. The model also supports dual-person mode, animating two speakers simultaneously within the same frame using separate audio tracks and bounding box annotations.

2. Key Features & Innovations

Triple Cross-Attention Audio Conditioning: Unlike text-only conditioned video models, InfiniteTalk injects audio embeddings at every transformer block via a dedicated cross-attention layer. Audio features are extracted frame-by-frame using a Wav2Vec2 encoder, providing per-frame speech signal anchoring that drives natural mouth movements, facial micro-expressions, and head motion synchronized to the audio input.

Streaming Long-Form Video Generation: InfiniteTalk's streaming mode processes audio in overlapping clip segments with configurable motion frame overlap, automatically concatenating segments into seamless long-form video. This enables generation of minutes-long talking videos without quality degradation or identity drift — a capability not available in standard image-to-video pipelines limited to single-shot outputs.

High-Fidelity Identity Preservation: The model maintains consistent facial identity, hairstyle, clothing texture, and background composition across the entire generated video. The audio conditioning signal provides strong per-frame constraints that prevent the identity drift commonly observed in long unconditional video generation.

Dual-Person Conversation Mode: InfiniteTalk supports animating two speakers in a single scene by accepting separate audio tracks and bounding box coordinates for each person. This enables realistic conversation scenarios, interview formats, and dialogue-driven content without requiring separate generation passes or post-production compositing.

Flexible Input Modalities: The model accepts either a static portrait image or a reference video as the visual source, combined with audio in WAV or MP3 format. Text prompts provide additional guidance for expression style, posture, and behavioral nuance, giving creators fine-grained control over the generated output.

Conditional VSR Upscaling: When generating at 720p resolution with audio duration under 60 seconds, InfiniteTalk automatically routes output through a FlashVSR super-resolution pipeline, delivering enhanced visual clarity without additional user configuration or cost management.

3. Model Architecture & Technical Details

InfiniteTalk is built on the Wan2.1 I2V-14B foundation model (14 billion parameters, 480p native resolution) with custom InfiniteTalk adapter weights that introduce the audio cross-attention pathway. The audio encoder uses a Chinese-Wav2Vec2-Base model that extracts frame-aligned speech embeddings at 25 fps video rate, creating a one-to-one correspondence between audio features and generated video frames.

The inference pipeline operates in two modes. In clip mode, the model generates a single video segment of up to 81 frames (approximately 3.2 seconds at 25 fps), suitable for short-form content. In streaming mode, the model iteratively generates overlapping clips with a configurable motion frame overlap (default: 9 frames), seamlessly blending segments to produce arbitrarily long video bounded only by the input audio duration and a configurable maximum frame limit.

The diffusion process uses a configurable number of denoising steps (default: 40, tunable from 1–100) with TeaCache acceleration for improved throughput. On NVIDIA H200 hardware, each 81-frame clip requires approximately 3.5 minutes of processing time, yielding a generation-to-output ratio of roughly 10–30× depending on resolution and hardware load.

For 720p output, the system employs a two-stage pipeline: base generation at 480p followed by conditional FlashVSR 4× upscaling (target: 921,600 pixels at 25 fps), applied automatically when audio duration is 60 seconds or less.

4. Performance Highlights

InfiniteTalk addresses a specific niche — audio-driven talking-head video — that differs from general-purpose text-to-video or image-to-video models. Its performance should be evaluated primarily on lip-sync accuracy, identity consistency, and long-form stability rather than visual diversity or cinematic motion range.

Capability	InfiniteTalk	General I2V Models	Dedicated Lip-Sync Tools
Lip-sync accuracy	Phoneme-level, multi-language	N/A (no audio input)	Word-level, often English-only
Maximum duration	Up to 10 minutes (streaming)	5–15 seconds typical	30–60 seconds typical
Identity preservation	High (audio-anchored per-frame)	Moderate (drift in longer clips)	Moderate
Dual-person support	Native	Not available	Rare
Resolution	480p native, 720p with VSR	Up to 1080p	Varies
Audio input	Any language WAV/MP3	N/A	Usually English TTS

InfiniteTalk achieves strong lip-sync fidelity across Chinese, English, Japanese, and other languages tested, owing to the language-agnostic Wav2Vec2 audio feature extraction. Identity drift is minimal even in 5+ minute generations due to the per-frame audio conditioning anchor.

5. Intended Use & Applications

Digital Avatar & Virtual Presenter: Create realistic talking-head videos for virtual hosts, AI assistants, and digital spokespersons using a single photo and recorded or synthesized speech audio.

Video Dubbing & Localization: Generate lip-synced video from translated audio tracks, enabling cost-effective multilingual content adaptation without re-filming or manual lip-sync editing.

Online Education & Training: Produce instructor-led video content at scale from lecture audio recordings and a single instructor photograph, reducing video production costs for e-learning platforms.

Podcast & Interview Visualization: Transform audio-only podcast or interview recordings into engaging video content with realistic speaker animations, suitable for social media distribution.

Customer Service & Chatbot Video: Generate personalized video responses driven by TTS audio output, enabling human-like video communication in automated customer interaction flows.

Social Media Content at Scale: Rapidly produce talking-head content for influencer accounts, news summaries, or commentary formats using text-to-speech pipelines combined with InfiniteTalk video generation.

Benzer Modelleri Keşfedin

NEW

Görüntü-Video

TURBO

Wan-2.2-turbo-spicy Image-to-video Lora

Fast image-to-video generation with custom LoRA support. Powered by Wan 2.2 rCM turbo with high/low noise LoRA injection. Supports 480p, 720p, and 1080p output.

Wan-2.2-turbo-spicy Image-to-video

Fast image-to-video generation powered by Wan 2.2 with rCM turbo acceleration. Supports 480p, 720p, and 1080p (via VSR upscaling) output with 5s or 8s duration.

Wan 2.2 Turbo Image-to-Video

Image-to-video model for fast single-clip generation with stable motion and 30fps workflow post-processing.

Wan 2.2 Turbo Infinite Image-to-Video

Image-to-video model for segmented prompt video generation with stable motion and 30fps workflow post-processing.

Wan 2.2 Turbo Infinite Image-to-Video LoRA

Image-to-video LoRA variant for segmented prompt video generation with stable motion and 30fps workflow post-processing.

Wan 2.2 Turbo Spicy Infinite Image-to-Video

Image-to-video model for segmented prompt video generation with stable motion and 30fps workflow post-processing.

Wan 2.2 Turbo Spicy Infinite Image-to-Video LoRA

Image-to-video LoRA variant for segmented prompt video generation with stable motion and 30fps workflow post-processing.

Video Upscaler

Upscale an existing video to 1080p or 2K while preserving motion, timing, and source composition. 4K support is planned for a later release.

Face Swap

Replace the person in a video with a face of your choice. Motion, timing and background are preserved while the character is swapped.

Van-2.6 Text-to-video

A speed-optimized text-to-video option that prioritizes lower latency while retaining strong visual fidelity. Ideal for iteration, batch generation, and prompt testing.

Van-2.6 Image-to-video

A speed-optimized image-to-video option that prioritizes lower latency while retaining strong visual fidelity. Ideal for iteration, batch generation, and prompt testing.

Wan 2.7 Spicy Image-to-Video

AtlasCloud Wan 2.7 Spicy Image-to-Video turns a first-frame image into short cinematic motion with stable temporal detail and expressive character movement.

Wan 2.6 Spicy Image-to-Video

AtlasCloud Wan 2.6 Spicy Image-to-Video turns a reference image into a short motion clip with expressive character movement and stable temporal detail.

Van-2.5 Image-to-video

Get animated visuals from your images faster without major quality sacrifice. Perfect for preview workflows, previews at scale, or mass production of animated assets.

Van-2.5 Text-to-video

Convert prompts into cinematic video clips with synchronized sound. Van 2.5 generates 720p/1080p outputs with stable motion, native audio sync, and prompt-faithful visual storytelling.

Sync.so Lipsync v3

Sync.so Lipsync v3 (sync-3) is Sync Labs state-of-the-art lip-sync model, re-syncing the lips of an existing video to a new audio track with industry-leading naturalness.

LIP_SYNC

From

$0.22/SANİYE

Tüm Medya Yapay Zekâsı için Tek API.

Tüm modelleri keşfet