Choose text or image
Start from text for a fresh composition or from an image when the first frame already works.
Create short fictional adult video scenes from a written direction or animate an existing generated frame. Choose the route, describe one cinematic beat and iterate with controlled changes.
Choose one visual beat, one movement and one camera idea.
LTX 2.3 supports two creative routes: text-to-video for building a scene from a written concept, and image-to-video when an existing fictional adult frame should anchor the composition and identity.
Short clips benefit from one visual beat. A clear motion verb, a simple camera instruction and a specific atmosphere give the model a readable sequence to solve. Iterate one variable at a time instead of rewriting every part of the direction.
Start from text for a fresh composition or from an image when the first frame already works.
Describe the subject, action, environment, pace and camera in a clear sequence.
Review identity, anatomy, direction of motion and continuity across the full result.
Adjust only the action, camera, timing or source frame so each change is measurable.
Use text when the composition is still open; use an image when the visual identity is already decided.
This choice changes how the prompt works. Text-to-video carries more responsibility for the subject and environment. Image-to-video can spend more of the instruction on movement, timing and the camera because the first frame already supplies the visual foundation.
Develop a new short scene directly from a compact cinematic description.
Animate an eligible generated frame while preserving its visual starting point.
Build the clip around one action and one readable emotional beat.
Compare focused changes to motion, timing, framing or atmosphere.
The same concept can be directed through subject motion, environmental motion or camera motion.
Give one type of movement priority. When every layer moves aggressively at the same time, the scene loses a stable visual anchor.
Define the fictional adult, location, action, atmosphere and camera from the prompt.
Let the image carry identity while the prompt concentrates on performance.
A slow push, tracking move or locked composition changes how the action feels.
Write the video prompt in the order the viewer experiences it.
Introduce the subject and environment, state the main action, describe its pace and end with the camera direction. For image-to-video, omit visual details already established by the source unless they must remain unchanged.
Original fictional adult character [description] in [setting]. The character [one action] at [pace] while [secondary natural motion]. [Lighting and atmosphere]. Camera [single movement or locked shot]. Clearly adult, no resemblance to a real person.Decide whether a written concept or a generated image is the better starting point.
Give the short clip a focused action that can be read from beginning to end.
Use one move that supports the action instead of competing with it.
Check the start, middle and final frame for identity, anatomy and visual continuity.
Start from text for discovery. Start from an image for stronger visual control.
If the character and composition are not settled, text-to-video can reveal new scene directions. If an existing image already has the right identity and framing, image-to-video lets you focus on how that exact moment changes.
Discovery needs space; continuity needs a stable visual anchor.
Quick answers about the LTX 2.3 workflow.
It is presented here as a short-form video model with text-to-video and image-to-video routes for original fictional adults.
Use text when you want the model to develop the composition. Use an image when identity, pose and framing are already established.
Long enough to define the subject, one action, pace, atmosphere and camera—without stacking unrelated events.
Begin from a clear generated source image, avoid contradictory instructions and keep movement within what the starting pose can support.
A short clip lets you evaluate one beat, make focused changes and avoid asking the model to maintain too many scene transitions at once.
No. Use original fictional adults only and never create deceptive or non-consensual sexual deepfakes.
Choose text or an image, define a focused cinematic beat and refine the result through controlled changes.
Try LTX 2.3 ↗