Hotel Lobby AI: choose a two-subject workflow
The orange booth is a visual starting point. Before writing a long prompt, decide what material you already have and what must remain recognizable. Different inputs solve different tasks. This guide prepares the handoff; it has not been validated with a paid render.
Written and documentation checked
1. Choose by the input you have
Only an idea? Start with an original scene draft. Have separate subject references? An ordered-reference workflow can express which input belongs to which person or pet. Have a finished composition? Describe motion instead of rebuilding its appearance in text. Have a source clip whose timing matters? Prepare a video-editing brief and confirm that the chosen service accepts that input.
| Path | Prepare | What the text does |
|---|---|---|
| Original scene | Two descriptions and a shared scene | Sets creative direction; supplies no visual identity evidence |
| Ordered references | Authorized left and right references | Maps inputs to distinct roles; check model-specific labels |
| Image to video | A finished first frame with both subjects | Describes subject and camera movement |
| Reference video edit | A cleared source clip and replacements | Records preservation and changes; does not extract motion |
2. Lock the left/right assignment before wording
Write a small reference map: left = the adult in a navy jacket; right = the ginger cat. Add local filenames in the helper so you can find the correct image later. In the documented Kling O1 reference interface, Element 1 and Element 2 refer to ordered elements. Other models may use different controls. Verify the actual interface instead of pasting those labels everywhere.
Changing sides is a data change, not merely a word change. Move the description, reference and individual action together. The helper’s Swap sides does this while keeping your edited draft intact until you regenerate.
3. Give each subject one observable action
For a finished first frame, Runway’s guidance separates the image’s visual information from the motion described in text. A useful first experiment is a small action for each side and one camera direction. This is an iteration method, not evidence that our draft will succeed.
Conflicting draft: “A locked camera circles both performers as they dance, change clothes and trade places.” Revised experiment: “The camera holds a two-shot. The left subject nods once. The right subject remains seated and turns its head slightly.” The second version makes it easier to identify what failed because fewer variables change. Neither example is a tested video case.
4. Separate preservation from replacement
For a source-video brief, name the clip and identify what should stay: subject placement, background, shot continuity or timing. Then specify each replacement. Do not promise a precise dance, lip sync or identity match from text alone. The provider’s video-input support and editing controls are part of the workflow; our text tool cannot inspect a clip.
5. Save a version, then inspect evidence
Name the draft “first motion test”, save it locally and export the production brief. At the provider, record the model/version, date, settings, input rights and actual result. Compare left/right identity, hands or paws, microphone position and camera movement. Change one direction at a time and keep failed results in your notes. The local history stores drafts, never videos.
Finished means you have a clear, reviewable handoff. Rendering, audio licensing and publication approval happen separately. No song, performance clip or claim of provider affiliation is included.
Sources and limits
Provider documentation supports the input distinctions. Practical checklists and example revisions here are original preparation advice. No real render, user study, success rate or cost result is claimed.