내레이션이 들어간 비디오 만들기

LLM으로 대본을 쓰고, 클립을 생성하고, 보이스오버를 합성해 하나로 합치는 엔드투엔드 과정 — API 키 하나로 끝냅니다.

이 튜토리얼은 하나의 스크립트 안에서 세 가지 모델 유형을 엮습니다. 언어 모델이 내레이션을 쓰고, 비디오 모델이 영상을 만들고, 음성 모델이 대본을 소리 내어 읽습니다. 하나의 API와 하나의 잔액으로 모델 유형을 넘나드는 것이 왜 유용한지 보여 주는 현실적인 예제입니다.

필요한 것: Python 3.9+, API 키, 그리고 마지막에 오디오와 비디오를 합치려면 ffmpeg.

이 튜토리얼은 실제 생성 작업을 실행하며 실제 크레딧을 소모합니다. 비디오가 가장 비싼 부분입니다 — 이것저것 시도하는 동안에는 짧은 길이로 시작하고, 정확한 금액을 미리 알고 싶다면 가격 추정을 이용하세요.

준비

pip install requests
export ATLASCLOUD_API_KEY="your-api-key"

1단계 — 공통 헬퍼

모든 미디어 작업은 제출한 뒤 폴링하는 동일한 형태를 따르므로, 한 번만 정의해 둡니다.

import os, time, requests

API_KEY = os.environ["ATLASCLOUD_API_KEY"]
BASE = "https://api.atlascloud.ai/api/v1"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
TERMINAL = {"completed", "succeeded", "failed", "timeout"}


def submit(endpoint: str, model: str, **params) -> str:
    """작업을 제출한다. 파라미터는 평평하게 두고 input으로 감싸지 않는다."""
    r = requests.post(f"{BASE}/model/{endpoint}",
                      headers=HEADERS, json={"model": model, **params}, timeout=60)
    r.raise_for_status()
    return r.json()["data"]["id"]


def wait(prediction_id: str, timeout: int = 900) -> dict:
    """종료 상태까지 폴링하며 간격을 점점 늘린다."""
    deadline, delay = time.time() + timeout, 2.0
    while time.time() < deadline:
        r = requests.get(f"{BASE}/model/prediction/{prediction_id}",
                         headers=HEADERS, timeout=30)
        r.raise_for_status()
        data = r.json()["data"]
        if data.get("status") in TERMINAL:
            if data["status"] in ("failed", "timeout"):
                raise RuntimeError(f"작업 실패: {data.get('error') or data['status']}")
            return data
        time.sleep(delay)
        delay = min(delay * 1.5, 10.0)
    raise TimeoutError(prediction_id)


def download(url: str, path: str) -> str:
    r = requests.get(url, timeout=300)
    r.raise_for_status()
    with open(path, "wb") as f:
        f.write(r.content)
    return path

2단계 — LLM으로 대본 쓰기

언어 모델은 기본 URL이 다르고 동기 방식이므로 submit/wait을 거치지 않습니다.

def write_narration(topic: str) -> str:
    r = requests.post(
        "https://api.atlascloud.ai/v1/chat/completions",
        headers={**HEADERS, "Content-Type": "application/json"},
        json={
            "model": "deepseek-ai/deepseek-v3.2",
            "messages": [
                {"role": "system",
                 "content": "You write narration for short videos. "
                            "Reply with two sentences of spoken narration and nothing else."},
                {"role": "user", "content": f"Topic: {topic}"},
            ],
            "max_tokens": 200,
        },
        timeout=120,
    )
    r.raise_for_status()
    return r.json()["choices"][0]["message"]["content"].strip()


narration = write_narration("how ocean waves shape a coastline")
print(narration)

3단계 — 영상 생성하기

video_id = submit(
    "generateVideo",
    "alibaba/wan-2.5/text-to-video",
    prompt="slow aerial shot of waves breaking against a rocky coastline at golden hour",
    duration=5,
)
video = wait(video_id)
download(video["outputs"][0], "footage.mp4")

비디오 생성은 초가 아니라 분 단위로 걸립니다. 프로덕션에서는 웹훅을 등록해, 폴링 루프를 계속 열어 두는 대신 작업이 끝나면 콜백을 받으세요.

4단계 — 보이스오버 합성하기

2단계에서 만든 내레이션이 여기서 입력이 됩니다. 음성, 음악, 전사는 모두 같은 엔드포인트를 공유하며, 어느 것이 될지는 모델이 결정합니다.

audio_id = submit(
    "generateAudio",
    "bytedance/seed-audio-1.0",
    text=narration,
    format="mp3",
    sample_rate=44100,
    speech_rate=-5,   # 조금 느린 편이 내레이션으로 알아듣기 좋다
)
audio = wait(audio_id)
download(audio["outputs"][0], "voiceover.mp3")

음성 레퍼런스를 포함한 전체 파라미터는 오디오 모델을 참조하세요.

5단계 — 합치기

ffmpeg -i footage.mp4 -i voiceover.mp3 \
  -c:v copy -c:a aac -shortest narrated.mp4

선택 사항 — 자막 생성하기

보이스오버를 전사 모델에 다시 통과시키면 자막용 단어 단위 타이밍을 얻을 수 있습니다.

stt_id = submit(
    "generateAudio",
    "bytedance/seed-asr-2.0",
    audio_url=audio["outputs"][0],   # 필드 이름이 audio_url임에 주의
    enable_punc=True,
    show_utterances=True,
)
stt = wait(stt_id)

# 전사 모델에서 outputs[0]은 파일 URL이 아니라 텍스트 자체다
print(stt["outputs"][0])
for word in stt.get("stt_result", {}).get("words", [])[:10]:
    print(word["start"], word["end"], word["text"])

전사 모델은 audio_url을 받지만, 일부 다른 음성 인식 모델은 대신 audio를 받습니다. 해당 모델의 API 레퍼런스를 확인하세요 — 필드 이름을 잘못 넘기면 검증에서 실패합니다.

다음으로 해 볼 것

  • 비디오 모델을 참조 이미지를 받는 모델로 바꾸고, 직접 제공한 스틸 이미지로 분위기를 결정해 보세요
  • 폴링을 웹훅으로 옮겨 긴 작업이 프로세스를 막지 않게 하세요
  • 여러 클립을 병렬로 실행한 뒤 이어 붙여 더 긴 시퀀스를 만드세요
  • 오류와 요청 제한의 재시도 처리를 추가하세요

관련 문서

Last updated on

On this page