ナレーション付き動画をつくる

LLM に台本を書かせ、クリップを生成し、ナレーションを合成して 1 本にまとめる — API キー 1 つで完結するエンドツーエンドの手順。

このチュートリアルでは、1 つのスクリプトの中で 3 種類のモデルをつなぎます。言語モデルがナレーションを書き、動画モデルが映像を生成し、音声モデルがその台本を読み上げます。API と残高がモデルの種類をまたいで 1 つにまとまっていることの利点が、実感できる例です。

必要なもの: Python 3.9+、API キー、そして最後に音声と映像を多重化(mux)したい場合は ffmpeg。

このチュートリアルは実際の生成ジョブを走らせ、実際のクレジットを消費します。動画がいちばん高くつく部分です — 試行錯誤の段階では短い長さから始めてください。事前に正確な金額を知りたい場合は料金の見積もりをご利用ください。

セットアップ

pip install requests
export ATLASCLOUD_API_KEY="your-api-key"

ステップ 1 — 共通ヘルパー

どのメディアジョブも「送信してからポーリング」という同じ形なので、一度だけ定義しておきます。

import os, time, requests

API_KEY = os.environ["ATLASCLOUD_API_KEY"]
BASE = "https://api.atlascloud.ai/api/v1"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
TERMINAL = {"completed", "succeeded", "failed", "timeout"}


def submit(endpoint: str, model: str, **params) -> str:
    """ジョブを送信する。パラメータはフラットに置き、input で包まない。"""
    r = requests.post(f"{BASE}/model/{endpoint}",
                      headers=HEADERS, json={"model": model, **params}, timeout=60)
    r.raise_for_status()
    return r.json()["data"]["id"]


def wait(prediction_id: str, timeout: int = 900) -> dict:
    """終了状態になるまでポーリングし、間隔を徐々に広げる。"""
    deadline, delay = time.time() + timeout, 2.0
    while time.time() < deadline:
        r = requests.get(f"{BASE}/model/prediction/{prediction_id}",
                         headers=HEADERS, timeout=30)
        r.raise_for_status()
        data = r.json()["data"]
        if data.get("status") in TERMINAL:
            if data["status"] in ("failed", "timeout"):
                raise RuntimeError(f"ジョブが失敗しました: {data.get('error') or data['status']}")
            return data
        time.sleep(delay)
        delay = min(delay * 1.5, 10.0)
    raise TimeoutError(prediction_id)


def download(url: str, path: str) -> str:
    r = requests.get(url, timeout=300)
    r.raise_for_status()
    with open(path, "wb") as f:
        f.write(r.content)
    return path

ステップ 2 — LLM で台本を書く

言語モデルは別のベース URL を使い、同期的に動作するため、submit / wait は通りません。

def write_narration(topic: str) -> str:
    r = requests.post(
        "https://api.atlascloud.ai/v1/chat/completions",
        headers={**HEADERS, "Content-Type": "application/json"},
        json={
            "model": "deepseek-ai/deepseek-v3.2",
            "messages": [
                {"role": "system",
                 "content": "You write narration for short videos. "
                            "Reply with two sentences of spoken narration and nothing else."},
                {"role": "user", "content": f"Topic: {topic}"},
            ],
            "max_tokens": 200,
        },
        timeout=120,
    )
    r.raise_for_status()
    return r.json()["choices"][0]["message"]["content"].strip()


narration = write_narration("how ocean waves shape a coastline")
print(narration)

ステップ 3 — 映像を生成する

video_id = submit(
    "generateVideo",
    "alibaba/wan-2.5/text-to-video",
    prompt="slow aerial shot of waves breaking against a rocky coastline at golden hour",
    duration=5,
)
video = wait(video_id)
download(video["outputs"][0], "footage.mp4")

動画生成には秒ではなく分単位の時間がかかります。本番環境では Webhook を登録し、ポーリングループを開いたままにするのではなく、ジョブ側からコールバックさせてください。

ステップ 4 — ナレーションを合成する

ステップ 2 のナレーションが、ここでの入力になります。音声合成・音楽・文字起こしはすべて同じエンドポイントを共有し、どれになるかはモデルが決めます。

audio_id = submit(
    "generateAudio",
    "bytedance/seed-audio-1.0",
    text=narration,
    format="mp3",
    sample_rate=44100,
    speech_rate=-5,   # 少し遅めのほうがナレーションとして聞き取りやすい
)
audio = wait(audio_id)
download(audio["outputs"][0], "voiceover.mp3")

ボイスリファレンスを含む全パラメータは音声モデルをご覧ください。

ステップ 5 — 合成する

ffmpeg -i footage.mp4 -i voiceover.mp3 \
  -c:v copy -c:a aac -shortest narrated.mp4

オプション — 字幕を生成する

ナレーション音声を文字起こしモデルに通すと、字幕用の単語単位のタイミングが得られます。

stt_id = submit(
    "generateAudio",
    "bytedance/seed-asr-2.0",
    audio_url=audio["outputs"][0],   # フィールド名は audio_url なので注意
    enable_punc=True,
    show_utterances=True,
)
stt = wait(stt_id)

# 文字起こしモデルでは outputs[0] はファイルの URL ではなくテキストそのもの
print(stt["outputs"][0])
for word in stt.get("stt_result", {}).get("words", [])[:10]:
    print(word["start"], word["end"], word["text"])

文字起こしモデルは audio_url を受け取りますが、音声認識モデルの中には代わりに audio を取るものもあります。モデルの API リファレンスを確認してください — フィールド名を間違えるとバリデーションで失敗します。

次のステップ

  • 動画モデルを参照画像に対応したものに差し替え、用意した静止画で見た目をコントロールする
  • ポーリングを Webhook に移し、長時間ジョブがプロセスをブロックしないようにする
  • 複数のクリップを並列に生成し、つなげて長いシーケンスにする
  • エラーとレート制限のリトライ処理を追加する

関連ドキュメント

Last updated on

On this page