How to Use AI Inpainting to Make 3D People Look Realistic in Architectural Renders
You can spend forty hours perfecting materials, lighting, and vegetation β and a single 3D person standing near the camera will still give the whole image away. Clients rarely say "the subsurface scattering on that figure is off." They say the render feels fake, and they can't tell you why.
The entourage problem is one of the most discussed pain points in architectural visualization, and every traditional way to fix 3D people in renders has a known failure mode: scanned 3D people look waxy up close, 2D cutouts never quite match your scene's lighting, and fully AI-generated people arrive with warped hands and smeared faces.
AI inpainting solves this differently: instead of placing a foreign asset into your scene, it regenerates just the person region so the figure inherits your render's actual lighting, color grade, and perspective. In this guide you'll learn a repeatable workflow to fix people in any render β in minutes, without touching your 3D scene or re-rendering a single frame.
Why People Are the Weakest Part of Most Renders
Humans are hardwired to read faces and bodies. We tolerate a slightly-off wood texture, but we detect an artificial person instantly β the uncanny valley applies at full force in archviz. Getting realistic people in architectural visualization is hard because the three standard approaches each break down in a predictable way:
- Scanned 3D people (AXYZ, Renderpeople, Human Alloy): excellent at mid and far distance, but near the camera the fixed pose, frozen expression, and simplified skin shading become obvious. Hair is usually the first giveaway β scan hair reads as a solid helmet.
- 2D cutout people: fast and light on render time, but the photo was taken under different lighting than your scene. A cutout shot on an overcast day dropped into a warm sunset render almost always floats, no matter how carefully you color-correct. Perspective mismatch between the photo's camera and yours makes it worse.
- AI-generated people composited in: modern image models produce convincing humans, but generating a person separately and pasting them in reinstates the same lighting-mismatch problem β and text-to-image people are notorious for six-fingered hands and distorted faces at small sizes.
The common thread: all three methods import a person made under different conditions. Inpainting flips the logic β it generates the person inside your image, conditioned on the pixels around them.
The Workflow: Fix People in 4 Steps
This workflow assumes you already have a finished render (or photo) with a problem person in it β a plastic-looking scan model, a floating cutout, or an AI figure with distorted features. You'll need an inpainting tool that accepts a mask and a text prompt; the steps below use VisioMake's Render Editor, but the principles apply to any inpainting workflow, including Stable Diffusion with ControlNet.
Step 1: Upload the render and mask the person
Load your full-resolution render and paint a mask over the entire figure β including their shadow and any reflections on glossy floors. The most common beginner mistake is masking only the body: if the old shadow survives, the new person won't match it, and the seam gives the edit away. Extend the mask 10β20 pixels beyond the figure's silhouette so the model can blend edges cleanly.
Step 2: Describe the person you want
Write a short, concrete prompt: who they are, what they're doing, what they wear. "A woman in a beige trench coat walking, seen from behind, natural motion" beats "realistic person." Two details matter most for archviz:
- Match the activity to the space. A lobby wants people walking with intent; a cafΓ© terrace wants people seated and talking. Wrong activity reads as staged even when the rendering is flawless.
- Prefer back and three-quarter views for near-camera figures. Faces attract scrutiny; a person seen from behind is both more believable and keeps the viewer's eye on the architecture.
Step 3: Generate and compare variations
Run 2β4 generations and compare them at 100% zoom. Check three things in order: contact shadows (does the figure sit on the ground plane?), light direction (are highlights on the same side as your sun or key light?), and edge quality (no halo or blur ring around the mask boundary). Because inpainting conditions on your image, lighting usually matches automatically β but contact shadows are worth verifying every time.
Step 4: Refine locally if needed
If a result is 90% right but a hand or face is off, don't regenerate the whole person β mask just the problem area and inpaint again with a tighter prompt ("natural relaxed hand holding a phone"). Small masks converge much faster than large ones, and iterating locally preserves the parts that already work.
Fixing the Three Classic Failure Cases
Case 1: The waxy near-camera scan person
Keep the figure's silhouette and pose β they're usually fine β and let inpainting re-skin them. This is how to use AI to enhance 3D scan people rather than replacing them: mask the person, prompt for the same demographic and clothing ("man in a navy suit walking, mid-stride"), and the model regenerates skin, hair, and fabric with photographic detail while the composition stays intact. This is the highest-value fix: near-camera people carry the most scrutiny.
Case 2: The floating 2D cutout
Cutouts fail on lighting and grounding, so mask generously β person, shadow area, and a margin of floor. Describe the person and their grounding ("woman standing on the polished concrete floor, soft shadow to the left"). The inpainted replacement is generated under your scene's light, so the classic pasted-on look all but disappears.
Case 3: The AI person with distorted hands or face
This is where local inpainting shines. Don't discard the image β mask only the broken hand or face (often under 5% of the figure) and regenerate with a specific prompt. Faces at small scale often just need one pass; hands sometimes take two or three. Simplifying the ask helps: a hand in a coat pocket or holding a coffee cup is far easier for a model to get right than open spread fingers.
Matching Light, Grain, and Depth of Field
Inpainting inherits most scene properties automatically, but three finishing checks separate a good edit from an invisible one:
- Grain and sharpness: if your render has film grain or a soft post-processed look, a tack-sharp inpainted region stands out. Apply your grain/sharpening pass after inpainting, not before, so it unifies the whole frame.
- Depth of field: a person inside your DOF falloff zone must share its blur. Either inpaint before applying DOF, or mention it in the prompt ("slightly soft focus, background figure").
- Color grade: heavy LUTs shift skin tones. As with grain, grade after inpainting when you can.
The rule of thumb: inpaint people as the last step before global post-production β after lighting and materials are final, before grain, grade, and vignettes.
| Approach | Realism up close | Lighting match | Time per figure | Typical cost | Best for |
|---|---|---|---|---|---|
| Scanned 3D people (Renderpeople, Chaos Anima) | Lowβmedium | Good (rendered in scene) | 10β30 min incl. re-render | ββ¬59 per posed model (more for rigged/animated), or library subscription | Mid/background crowds |
| 2D cutout people | Medium | Poor (baked-in photo lighting) | 15β45 min of compositing | Freeββ¬15 per cutout | Tight deadlines, far background |
| Separately AI-generated + composited | Mediumβhigh | Poorβmedium | 20β40 min | Varies | Legacy workflow; rarely needed |
| AI inpainting in the render | High | Excellent (conditioned on your pixels) | 2β5 min, no re-render | ββ¬0.10 (10 credits) per generation | Near-camera and hero figures |
When Inpainting Is the Wrong Tool
Honesty about limits builds better workflows. Skip inpainting when:
- You need the same person across many frames. Inpainting generates a new person per image; an animation or a multi-view set of stills needs a consistent 3D figure. Use inpainting for hero stills and scan people for the animation.
- The figure must match a real individual β a client's team photo on the office render, for example. That's a compositing job, not a generation job.
- Dozens of small background people. Distant scan people or cutouts already look fine at 40+ pixels tall; inpainting each one is wasted effort. Fix the two or three figures nearest the camera and leave the crowd alone.
The 80/20 in practice: inpaint the foreground, populate the background conventionally.
Edit Any Part of Your Render Without Starting Over
Add reference images from your library β specific furniture, materials, or objects β place them directly on your render, and let AI blend them in naturally. Or mask any area and describe what should appear instead. Either way: seamless edits in seconds, no re-render needed.
Try it now