Hotel Lobby AI: How to Make the Trend Video

Learn what Hotel Lobby AI means and how to make your own short video with character references, templates, prompts, and AI video tools.

By VioEvo Editorial•Published 1 oktober 2026•Reading time 10 min

Tags

ai-video
image-to-video
reference-to-video
video-trends

Try a reference-driven scene before you read on.

The Hotel Lobby AI trend turns a familiar performance into a reusable video template. Creators keep the recognizable staging and rhythm, then swap in new characters, outfits, and visual styles with AI video tools.

This guide explains what people mean by Hotel Lobby AI, why the format spreads quickly, and how to make a short version with a consistent pair of characters. The workflow uses reference images and a short motion brief, so you can develop an original scene without downloading or re-uploading someone else's video.

What Is the Hotel Lobby AI Trend?

The trend uses the visual language of “Hotel Lobby,” a song by Quavo and Takeoff. One widely referenced source is the official COLORS performance:

Source: Quavo & Takeoff — HOTEL LOBBY | A COLORS SHOW

In the AI version, a creator usually prepares two character references, places them into a similar performance setup, and generates a new clip with a video model. The characters might be fictional, the creator's own avatars, or public figures. The format is recognizable because the template stays familiar while the people in it change.

You may also see the broader phrase hotel lobby video. It can refer to the original performance, AI trend tutorials, template videos, or ordinary hotel-lobby footage. Hotel Lobby AI usually means the how-to workflow for making a version of the trend.

Why These Videos Spread So Quickly

The format gives short-form creators a useful combination of structure and variation:

  • The source performance supplies a recognizable setting, pace, and visual idea.
  • A new pair of characters creates an immediate joke or surprise.
  • Character sheets make the result easier to repeat than a single improvised selfie.
  • A short clip can be tested, replaced, and shared without building a complete story.

That structure also explains why the trend is useful as an AI video exercise. It tests the parts of a workflow that normally cause trouble: preserving identity, positioning two subjects, directing a simple camera move, and keeping the scene coherent across several attempts.

What You Need Before You Generate

Prepare these elements first:

  1. Two character references. Use clear images with visible faces, stable lighting, and enough of the clothing or silhouette to distinguish each subject.
  2. A visual relationship. Decide who stands on the left and right, who is closer to the camera, and which details must remain unchanged.
  3. One short action. A confident step toward camera, a synchronized gesture, or a small performance beat is easier to control than a full dance routine.
  4. A frame plan. Vertical 9:16 works well for Shorts and Reels. Decide whether you want a medium shot, a wide lobby reveal, or a closer performance frame.
  5. Materials you can use. Use characters, music, and source footage that you have permission to use. You can create an original Hotel Lobby-style scene without using the original audio or copying the original performance shot for shot.

The Hotel Lobby AI Workflow

Step 1: Build Consistent Character Sheets

A character sheet gives the video model more than a name or a single face. Include a front view, a three-quarter view if available, the intended outfit, and one or two identity-bearing details. Keep the visual treatment consistent across both characters.

Useful details include:

  • hairstyle and hair color
  • jacket, shirt, or accessory that should stay fixed
  • approximate age range and body shape
  • expression and attitude
  • the side of the frame where each character begins

If you use real people, obtain permission for their likeness. For a safer example, use fictional characters or people whose images you control.

Step 2: Choose a Reference or Template Path

There are three practical ways to direct the scene:

  • Reference to video: provide character images and describe the new scene. This is the best starting point when identity and composition matter.
  • Image to video: animate one prepared key image when the scene is simple and the first frame is already close to the result you want.
  • Video to video: use a motion reference that you own or have permission to transform, then describe how the subjects and setting should change.

For a two-character trend video, start with VioEvo's reference-to-video workflow. It lets the references establish who is in the scene while the prompt controls the action, camera, and atmosphere.

Step 3: Replace the Characters With a Focused Prompt

Write the prompt as a shot brief. State what must remain stable before describing the movement:

Two consistent fictional characters stand in a polished hotel lobby at night. Character A remains on the left and Character B remains on the right. Keep their faces, hairstyles, clothing, and body proportions stable. They perform one confident, rhythmic movement toward the camera. Use a medium-wide vertical frame with a slow forward camera move, warm lobby lighting, a reflective marble floor, and restrained background motion.

The prompt does not need to mention that the result should be “viral.” The model needs concrete visual instructions: identity, position, one action, camera movement, and a small number of environmental details.

Step 4: Generate Short Variations

Make several short attempts and change one variable at a time. For example:

  • keep the characters fixed and change only the camera move;
  • keep the camera fixed and change only the gesture;
  • keep the action fixed and compare a realistic lobby with a stylized one;
  • keep the scene fixed and test a faster model before making a final quality pass.

This makes failures easier to diagnose. If the face drifts, strengthen the reference. If the action feels chaotic, reduce it to one visible change. If the camera invents a new shot, specify one camera movement and what should remain stable.

Prompt Variations You Can Adapt

Realistic Music-Video Version

Two fictional performers with consistent faces and clothing stand in a modern hotel lobby at night. They exchange one synchronized hand gesture and take a small step toward the camera. Medium-wide vertical composition, slow forward camera movement, warm practical lights, reflective marble floor, shallow background activity, stable identity and wardrobe, no sudden cuts.

Stylized Animation Version

Two original animated characters perform a short synchronized gesture in a brightly lit hotel lobby. Keep the character designs, colors, and positions consistent. Use a gentle push-in, clean graphic shapes, controlled reflections, and a playful performance energy. One action only, vertical short-form framing.

Fictional-Character Version

Two original fictional characters, one in a red varsity jacket and one in a cream trench coat, stand on opposite sides of a luxury hotel lobby. They turn toward each other, then face the camera together. Keep the outfits and facial features unchanged. Slow dolly in, polished stone floor, soft golden lighting, subtle background movement, vertical video.

Which VioEvo Workflow Should You Use?

Your starting pointRecommended workflow
You already have character imagesImage to video
You need several images to guide one sceneReference to video
You own a motion clip and want to transform itVideo to video
You are testing many ideas quicklyStart with a Fast model, then keep the strongest take
You need a more finished final passCompare a Quality model after the composition works

The model choice should follow the problem you are solving. Faster models are useful while you are deciding on the characters and movement. A higher-quality pass makes more sense after the shot is already working. See the Veo 3.1 model guide and Seedance 2.5 model guide for broader model context.

Common Problems and Quick Fixes

ProblemWhat to change
One character's face changes between framesUse a clearer reference and repeat the identity details once in the prompt
The two characters swap sidesState left and right positions and keep the camera mostly frontal
Clothing changes during the actionName the specific garment and add a stability instruction
The movement becomes a full dance sequenceReduce the clip to one gesture or one step
The camera invents dramatic movementsSpecify one camera move and a fixed shot size
The hotel lobby warps in the backgroundSimplify the architecture and keep background motion restrained
The result feels like a copy of the sourceChange the characters, setting details, action, and audio rather than reproducing the original shot exactly

The most efficient rewrite is usually small. Identify the first failure a viewer notices, then change the instruction connected to that failure instead of replacing the entire prompt.

Using Third-Party Videos and Real People Responsibly

The official performance and its music are copyrighted. The player above lets you watch the original YouTube upload in context. If you make your own version, use music and footage you have permission to use; watching an embedded video does not give you permission to download, edit, or reuse its underlying audio.

This AI parody is included as a reference example:

Source: When Elon Musk and Sam Altman start singing “Hotel Lobby”... by Dotiki AI.

This example helps explain the format. It does not show that the people depicted endorsed the video or VioEvo. Using a public figure's name, face, voice, or likeness can create additional rights and platform-rule issues, especially when a synthetic video could be mistaken for a real statement or endorsement. Use permission-based or fictional characters when you make your own version, and label synthetic media clearly when context could be misunderstood.

Music, footage, and character images may each have separate rights. If you plan to publish a new edit, check the permissions for each element first. YouTube's fair use guidance explains why an exception should not be assumed automatically.

Is Hotel Lobby AI Worth Trying?

Yes, if you treat it as a compact production exercise rather than a promise of guaranteed reach. The trend may fade quickly, but the skills behind it are durable: character consistency, reference-driven generation, restrained motion, and short-form iteration.

The best version for a brand, creator, or product team is usually an original scene that borrows the idea of a repeatable performance template while changing the characters, setting details, action, and audio. That gives the video a clear connection to the trend without making the result dependent on copying someone else's work.

Frequently Asked Questions

What is Hotel Lobby AI?

Hotel Lobby AI is a short-form video trend built around the “Hotel Lobby” performance by Quavo and Takeoff. Creators use AI to place new characters into a similar performance concept, often with character references and a video-generation model.

How do I make a Hotel Lobby AI video?

Prepare consistent character images, choose a reference-driven workflow, describe one short action and one camera move, then generate several variations. Start with reference to video when multiple visual references need to guide the scene.

Can I use any person in the video?

You should have permission to use a person's image, voice, or likeness. For public figures, avoid creating a video that could be mistaken for an authentic statement or endorsement.

Which AI video model works best for this trend?

There is no single best model for every attempt. Use a fast option while testing composition and motion, then compare a higher-quality model once the references and prompt are stable. The active model options are listed in the VioEvo generator.

Do I need the original “Hotel Lobby” song?

No. You can make an original performance-style video with rights-cleared audio, your own sound design, or no music during generation. Do not assume that a public upload gives you permission to reuse the song in a new commercial edit.

Can I make the video without copying the original performance?

Yes. Keep the general idea of a repeatable two-character performance, then create a different setting, movement, camera plan, and audio direction. This is also easier to adapt to your own characters and brand.