Hotel Lobby AI Video Generator

Make a Hotel Lobby AI video from two photos, people or pets: a duet at one hanging mic in an orange room. 10 seconds costs 125 credits. No song included.

  • From 63 credits per clip
  • From your photos
  • 5s / 10s / 15s
  • 480P
  • Audio available
Model

MiniMax H3 Max

Browse by task

Full or half body, face visible, one subject per photo. People who agreed, or your pets. No celebrities. People must be adults; no photos of children. JPG, PNG or WebP, up to 10 MB each.

What scene would you like to create?

A scene fills the prompt below. It calls your photos Image 1, Image 2 and so on; keep those names when you edit it.

Prompt0/2000
More settings · 9:16 · 10s · Audio
Aspect ratio
Duration
Resolution
Audio

Checking generation and private storage configuration.

Billing and delivery

Generate charges the displayed quote for the current length and resolution. If generation fails or the file fails our check, the credits come back automatically; a finished clip you don’t like is still charged. Finished clips stay in My videos.

Estimated cost: 125 credits
Sign in to generate

Sample result

10s · MP4
Two people, one hanging microphone.Made on Cicadas · MiniMax H3 Max · 10s
Made from these photos (AI-made, fictional)
  • Left sideLeft side
  • Right sideRight side
Use this recipe

About Hotel Lobby AI on Cicadas

Hotel Lobby AI puts two subjects, people or pets, on either side of one hanging microphone in a seamless orange room, trading lines in a locked-off full-body shot. Cicadas makes it from two photos with MiniMax H3 Max as a new vertical clip of 5, 10 or 15 seconds. The song from the trend is not included, and nothing is swapped onto the original performance. A 10-second clip is 125 credits.

Hotel Lobby AI examples made on Cicadas

From your photos10s · 768P

Two people, one hanging microphone.

MiniMax H3 Max, 10 seconds with sound at 768P, vertical, from two photos. A locked-off full-body shot: both faces and outfits match the photos as they trade lines. A studio ceiling with spotlights shows above the orange wall, and the model added a rapped voice in made-up words although the prompt asked for no vocals.

Made from these photos (AI-made, fictional)
  • Left sideLeft side
  • Right sideRight side
Use this recipe ↗
From your photos10s · 480P

A man and his cat at the mic.

MiniMax H3 Max, 10 seconds with sound, vertical, from a photo of a man and a photo of a tabby cat. The cat sits on a stool and looks up at the microphone with its mouth open while he raps beside it; its markings match the photo. The sound is a hip-hop beat with a voice in made-up words.

Made from these photos (AI-made, fictional)
  • Left sideLeft side
  • Right sideRight side
Use this recipe ↗

Why use Cicadas for Hotel Lobby AI

Left and right: two photos, people or pets

The first photo stands on the left, the second on the right; one click swaps them. A pet sits on a stool at microphone height. In our recorded clips both faces, the outfits and the cat's markings matched the photos.

The Hotel Lobby AI look, without the song

One hanging microphone, an orange room, a locked-off full-body shot. MiniMax H3 Max generates its own sound: in both recorded clips that was a hip-hop beat with a rapped voice in made-up words, although the prompt asked for no vocals. Replace it with your own sound when you post.

See the price of every Hotel Lobby AI video first

At 480P, 5 seconds is 63 credits, 10 seconds is 125 and 15 seconds is 188; at 768P they are 100, 200 and 300. A failed generation returns its credits.

How to use Hotel Lobby AI

  1. 01

    Sign in and add two photos, one for each side: full or half body, face visible.

  2. 02

    Pick a scene (two people, you and your pet, or two pets), keep 10 seconds, and check the quote.

  3. 03

    Tick the photo confirmation, generate, then download the MP4 from My videos and add your own sound when you post.

Hotel Lobby AI prompt ideas

Copy one into the prompt box above and adapt it to your shot.

  • “Vertical 9:16, one locked-off wide shot in a seamless bright orange studio, both people in frame from head to shoes for the whole clip. The person from Image 1 stands on the left and the person from Image 2 stands on the right, facing each other across one silver microphone that hangs from the ceiling between them. They take turns leaning toward the microphone and rapping with relaxed hand gestures, then nod along together. Keep both faces, hair and clothes exactly as in the photos. The camera never moves or cuts. Even studio light. Sound: a laid-back hip-hop drum beat, no vocals. No text, no logos.”

  • “Vertical 9:16, one locked-off full-body shot in a seamless bright orange studio. The person from Image 1 stands on the left. On the right, the pet from Image 2 sits upright on a tall wooden stool, level with one silver microphone that hangs from the ceiling between them. The person leans toward the microphone and raps with relaxed hand gestures; then the pet leans in, opens its mouth and bobs its head as if rapping back. Keep the person and the pet exactly as in the photos. Even studio light. Sound: a laid-back hip-hop drum beat, no vocals. No text, no logos.”

  • “Vertical 9:16, one locked-off full-body shot in a seamless bright orange studio. The pet from Image 1 sits upright on a tall wooden stool on the left and the pet from Image 2 sits on a matching stool on the right, facing each other across one silver microphone that hangs from the ceiling between them. They take turns leaning toward the microphone, opening their mouths and bobbing their heads as if rapping. Keep both pets exactly as in the photos. Even studio light. Sound: a laid-back hip-hop drum beat, no vocals. No text, no logos.”

Hotel Lobby AI FAQ

Tested on Cicadas · reviewed 2026-10-05

Recorded on 2026-10-05 with MiniMax H3 Max reference-to-video through fal.ai: a 10-second two-person clip at 768P (200 credits here) and a 10-second clip of a man and a cat at 480P (125 credits), both 9:16 with sound. The provider returned them in 19 and 8 seconds. Faces, outfits and the cat's markings matched the photos. Both soundtracks were a hip-hop beat with a voice in unintelligible words (two automated listeners), although the prompts asked for no vocals. A first two-person run at 768P came out cropped at the thighs; its wording was changed and it is not shown. The two-pets scene, 5 and 15 seconds and group photos were not run.

Updated Oct 5, 2026 · How we test