Description

Get extended video clips ready to post, without touching a single slider. Upload your images (JPG, PNG, WEBP, HEIC, up to 200MB), say something like "turn th...

README (SKILL.md)

Getting Started

Ready when you are. Drop your images here or describe what you want to make.

Try saying:

"convert a single landscape photo or a set of product images into a 1080p MP4"
"turn this image into a 2-minute video with smooth motion and background music"
"converting still images into long-duration animated videos for marketers, content creators, social media managers"

First-Time Connection

When a user first opens this skill, connect to the processing backend automatically. Briefly let them know (e.g. "Setting up...").

Authentication: Check if NEMO_TOKEN is set in the environment. If it is, skip to step 2.

Obtain a free token: Generate a random UUID as client identifier. POST to https://mega-api-prod.nemovideo.ai/api/auth/anonymous-token with header X-Client-Id set to that UUID. The response data.token is your NEMO_TOKEN — 100 free credits, valid 7 days.
Create a session: POST to https://mega-api-prod.nemovideo.ai/api/tasks/me/with-session/nemo_agent with Authorization: Bearer \x3Ctoken>, Content-Type: application/json, and body {"task_name":"project","language":"\x3Cdetected>"}. Store the returned session_id for all subsequent requests.

Keep setup communication brief. Don't display raw API responses or token values to the user.

Image to Video Long Duration — Convert Images into Extended Videos

Name: Image To Video Long Duration
Author: whitejohnk-26

This tool takes your images and runs long video creation through a cloud rendering pipeline. You upload, describe what you want, and download the result.

Say you have a single landscape photo or a set of product images and want to turn this image into a 2-minute video with smooth motion and background music — the backend processes it in about 1-3 minutes and hands you a 1080p MP4.

Tip: higher contrast images with clear subjects produce smoother and more realistic long-duration motion.

Matching Input to Actions

User prompts referencing image to video long duration, aspect ratio, text overlays, or audio tracks get routed to the corresponding action via keyword and intent classification.

User says...	Action	Skip SSE?
"export" / "导出" / "download" / "send me the video"	→ §3.5 Export	✅
"credits" / "积分" / "balance" / "余额"	→ §3.3 Credits	✅
"status" / "状态" / "show tracks"	→ §3.4 State	✅
"upload" / "上传" / user sends file	→ §3.2 Upload	✅
Everything else (generate, edit, add BGM…)	→ §3.1 SSE	❌

Cloud Render Pipeline Details

Each export job queues on a cloud GPU node that composites video layers, applies platform-spec compression (H.264, up to 1080x1920), and returns a download URL within 30-90 seconds. The session token carries render job IDs, so closing the tab before completion orphans the job.

All calls go to https://mega-api-prod.nemovideo.ai. The main endpoints:

Session — POST /api/tasks/me/with-session/nemo_agent with {"task_name":"project","language":"\x3Clang>"}. Gives you a session_id.
Chat (SSE) — POST /run_sse with session_id and your message in new_message.parts[0].text. Set Accept: text/event-stream. Up to 15 min.
Upload — POST /api/upload-video/nemo_agent/me/\x3Csid> — multipart file or JSON with URLs.
Credits — GET /api/credits/balance/simple — returns available, frozen, total.
State — GET /api/state/nemo_agent/me/\x3Csid>/latest — current draft and media info.
Export — POST /api/render/proxy/lambda with render ID and draft JSON. Poll GET /api/render/proxy/lambda/\x3Cid> every 30s for completed status and download URL.

Formats: mp4, mov, avi, webm, mkv, jpg, png, gif, webp, mp3, wav, m4a, aac.

Headers are derived from this file's YAML frontmatter. X-Skill-Source is image-to-video-long-duration, X-Skill-Version comes from the version field, and X-Skill-Platform is detected from the install path (~/.clawhub/ = clawhub, ~/.cursor/skills/ = cursor, otherwise unknown).

Include Authorization: Bearer \x3CNEMO_TOKEN> and all attribution headers on every request — omitting them triggers a 402 on export.

Draft JSON uses short keys: t for tracks, tt for track type (0=video, 1=audio, 7=text), sg for segments, d for duration in ms, m for metadata.

Example timeline summary:

Timeline (3 tracks): 1. Video: city timelapse (0-10s) 2. BGM: Lo-fi (0-10s, 35%) 3. Title: "Urban Dreams" (0-3s)

Backend Response Translation

The backend assumes a GUI exists. Translate these into API actions:

Backend says	You do
"click [button]" / "点击"	Execute via API
"open [panel]" / "打开"	Query session state
"drag/drop" / "拖拽"	Send edit via SSE
"preview in timeline"	Show track summary
"Export button" / "导出"	Execute export workflow

Reading the SSE Stream

Text events go straight to the user (after GUI translation). Tool calls stay internal. Heartbeats and empty data: lines mean the backend is still working — show "⏳ Still working..." every 2 minutes.

About 30% of edit operations close the stream without any text. When that happens, poll /api/state to confirm the timeline changed, then tell the user what was updated.

Error Handling

Code	Meaning	Action
0	Success	Continue
1001	Bad/expired token	Re-auth via anonymous-token (tokens expire after 7 days)
1002	Session not found	New session §3.0
2001	No credits	Anonymous: show registration URL with `?bind=\x3Cid>` (get `\x3Cid>` from create-session or state response when needed). Registered: "Top up credits in your account"
4001	Unsupported file	Show supported formats
4002	File too large	Suggest compress/trim
400	Missing X-Client-Id	Generate Client-Id and retry (see §1)
402	Free plan export blocked	Subscription tier issue, NOT credits. "Register or upgrade your plan to unlock export."
429	Rate limit (1 token/client/7 days)	Retry in 30s once

Common Workflows

Quick edit: Upload → "turn this image into a 2-minute video with smooth motion and background music" → Download MP4. Takes 1-3 minutes for a 30-second clip.

Batch style: Upload multiple files in one session. Process them one by one with different instructions. Each gets its own render.

Iterative: Start with a rough cut, preview the result, then refine. The session keeps your timeline state so you can keep tweaking.

Tips and Tricks

The backend processes faster when you're specific. Instead of "make it look better", try "turn this image into a 2-minute video with smooth motion and background music" — concrete instructions get better results.

Max file size is 200MB. Stick to JPG, PNG, WEBP, HEIC for the smoothest experience.

Export as MP4 with H.264 codec for the best balance of file size and playback compatibility.

Usage Guidance

This skill appears to do what it says: it uploads images to a nemovideo.ai backend, creates render jobs, and returns downloadable MP4s. Before installing, consider: (1) Privacy — your images are uploaded to https://mega-api-prod.nemovideo.ai; avoid uploading sensitive images unless you trust the service and its retention/policy. (2) Token handling — the skill will use NEMO_TOKEN (or obtain an anonymous token for 7 days if none is provided); treat that token like a credential and revoke it if you stop using the service. (3) Transparency — SKILL.md says not to show raw API responses or tokens to users; verify that you’re comfortable with the skill making network calls on your behalf. (4) Source verification — there is no homepage or known publisher listed; if you need higher assurance, verify the vendor/endpoint and check their privacy and terms before use. Otherwise, the skill’s requested access and instructions are proportionate to its stated purpose.

Capability Analysis

Type: OpenClaw Skill Name: image-to-video-long-duration Version: 1.0.0 The skill provides a functional integration for the NemoVideo image-to-video service. It manages authentication via environment variables or anonymous token generation, handles session-based cloud rendering, and processes file uploads to 'mega-api-prod.nemovideo.ai'. The instructions in SKILL.md are well-defined for the AI agent to manage API states and error codes without any evidence of data exfiltration, unauthorized access, or malicious intent.

Capability Assessment

✓ Purpose & Capability

Name/description, required env var (NEMO_TOKEN), declared config path (~/.config/nemovideo/), and the documented API endpoints all align: the skill is a cloud render frontend for nemovideo.ai and does not request unrelated services or credentials.

ℹ Instruction Scope

SKILL.md directs the agent to obtain/store an anonymous token if NEMO_TOKEN is absent, create a session_id, upload user image files, open SSE chat streams, poll render status, and return download URLs. These actions are within the expected scope. The instructions also suggest detecting install platform by checking common install paths (~/.clawhub/ or ~/.cursor/skills/) to set an X-Skill-Platform header; this requires reading the agent install path (minor scope expansion) but is not necessary for core functionality.

✓ Install Mechanism

No install spec or code files are present (instruction-only), so nothing is written to disk by the skill itself and there are no third-party install URLs to evaluate.

✓ Credentials

Only a single credential (NEMO_TOKEN) is required and is the declared primaryEnv. That token is justified by the skill's need to authenticate to the nemovideo.ai API. No unrelated secrets or broad environment access are requested.

✓ Persistence & Privilege

always is false and the skill does not request permanent system-wide changes. It instructs storing a session_id and using/optionally persisting a NEMO_TOKEN for subsequent calls (normal for a session-based cloud API). It does not modify other skills or system settings.

Version History

v1.0.0

Initial release of Image to Video Long Duration - Convert uploaded images (JPG, PNG, WEBP, HEIC) into extended 1080p videos with motion and background music. - Automatic backend session connection and authentication with free credits for new users. - Simple upload & command workflow: describe your desired video style, receive downloadable MP4 results. - Session-based timeline supports advanced editing, batching, and iterative workflows. - Export, credit balance, and state management handled via dedicated cloud endpoints. - Comprehensive error handling for invalid tokens, unsupported formats, and export restrictions.

Metadata

Slug image-to-video-long-duration

Version 1.0.0

License MIT-0

All-time Installs 0

Active Installs 0

Total Versions 1

Frequently Asked Questions

What is Image To Video Long Duration?

Get extended video clips ready to post, without touching a single slider. Upload your images (JPG, PNG, WEBP, HEIC, up to 200MB), say something like "turn th... It is an AI Agent Skill for Claude Code / OpenClaw, with 64 downloads so far.

How do I install Image To Video Long Duration?

Run "/install image-to-video-long-duration" in the OpenClaw or Claude Code chat to install it in one step — no extra setup required.

Is Image To Video Long Duration free?

Yes, Image To Video Long Duration is completely free, licensed under MIT-0. You can download, install and use it at no cost.

Which platforms does Image To Video Long Duration support?

Image To Video Long Duration is cross-platform and runs anywhere OpenClaw / Claude Code is available (cross-platform).

Who created Image To Video Long Duration?

It is built and maintained by whitejohnk-26 (@whitejohnk-26); the current version is v1.0.0.

More Skills

Image To Video Long Duration