Chapter 03 · Section II · 16 min read
Image generation for Nepali audiences
Current image models default to a generic South-Asian aesthetic that reads Indian, not Nepali — and the honest fix is heavy specification, careful review, and knowing when to stop and shoot real photography.
Type “a happy Nepali family at Dashain” into Midjourney, DALL·E, or the Adobe generator and watch what comes back. A woman in a saree that drapes Mumbai. A man with a beard and a kurta that reads Lucknow. A backdrop of marigolds and an oil lamp that is either Diwali furniture or a generic temple courtyard from somewhere in Uttar Pradesh. The image is fluent. It is also visibly, embarrassingly, not Nepal. This is the single most important practical issue with image generation for a Nepali marketing team in 2026, and it is not going to be fixed by waiting for a better model. The fix is in the brief, the prompt, the review process, and — most importantly — knowing when to put the model down and call a photographer.
The default-Indian problem
The training data tells you what is going to happen before you type. Image models are trained on hundreds of millions of captioned images scraped from the web. The web’s South-Asian content is overwhelmingly Indian — Bollywood stills, Indian wedding photography, Indian tourism marketing, Indian editorial fashion. Nepali content on the indexed web is a thin slice in comparison, and the model has been given no strong reason to treat Nepal as distinct from “South Asia” in general. When you type “Nepali”, the model averages across what it has seen labelled South Asian and gives you back the mode of that distribution. The mode is Indian.
The result is a consistent set of failures. The “Nepali” bride wears a heavy red Indian-style lehenga, not a Newari hakupatasi or a Tarai gown. The “Nepali” man wears a beige kurta, not a daura suruwal with a dhaka topi. The “Nepali” family sits in a Mumbai-coded living room, not a Kathmandu Valley courtyard with carved windows and a household shrine. The “Nepali” mountain in the background looks like a generic Himalayan peak from a stock library, not a recognisable one from this country. None of this is malicious. It is what averaging across an Indian-dominated dataset produces when you give the model a one-word ethnic label and ask it to fill in the rest.
This matters because your audience is not averaging. A Nepali viewer reads the saree-and-Mumbai-living-room image as “this brand thinks I am Indian,” which is the single fastest way to lose trust on the platform. The image does not have to be wrong in a way you can articulate. It has to land wrong in a way the audience feels immediately.
Heavy-specific prompting as the partial fix
The model is correctable. Not perfectly, not always, but enough to be useful — provided you stop using one-word ethnic labels and start giving it the specifics a Nepali photographer would carry in their head before a shoot.
Specific ethnic and regional cues do most of the work. “A Newari household in Patan during Dashain” lands closer to the truth than “a Nepali household.” “A Limbu wedding ceremony in Taplejung” gives the model entirely different material to draw from than “a Nepali wedding.” “A Tharu family in a Dang village courtyard during Maghi” is concrete in a way “a Madhesi family” is not. The model has seen at least some images labelled with these specific terms; the more specific your cue, the more likely the model pulls from the right cluster.
Specific dress instructions matter more than anything else in the visual frame. “Daura suruwal with a black dhaka topi for the older men; gunyo cholo for the older women; modern-cut kurtha-suruwal for the younger women; the children in school-uniform-like clothes.” That sentence does more for the output than a paragraph of mood. Without it the model defaults to the nearest South-Asian wardrobe it has seen most often.
Specific backdrop instructions anchor the geography. “Carved wooden windows of a traditional Patan home, brick courtyard floor, a brass household shrine in one corner, marigold strings tied across the doorway, soft late-afternoon light from the courtyard sky” tells the model where you actually are. “A Phewa Lake shoreline at sunrise with the Annapurna range visible in the distance, fishing boats moored in the foreground” is a specific Pokhara morning, not a generic lake. “A Terai rice paddy in mid-Asar, ankle-deep water, planting in progress, a small mud-brick farmhouse at the field edge” is a Madhesh scene the model can compose toward if you ask for it precisely.
Specific colour palette closes the loop. “Warm earth tones, deep reds, brass, marigold orange, soft natural light — no neon, no saturated Bollywood pinks, no Holi-coloured powder.” The avoid clauses in the colour brief do as much work as the positive ones.
The brand-safety risks specific to AI people
Even when you get the cultural specifics right, AI-generated humans carry a separate set of risks that a reviewer has to look for before anything ships.
Hands and fingers. The single most reliable tell. Six fingers, three knuckles in the wrong place, a thumb attached at the wrong angle, a hand fused into a sleeve. Close-ups on hands almost always need a regenerate or a Photoshop fix. Holding a tika plate, performing namaste, exchanging money — anywhere hands are doing work, look closely.
Eyes. Slightly mismatched pupils, irises that drift different directions, a glassy quality on the second person in a two-person scene who the model paid less attention to. A face that reads warm on a casual glance and uncanny on a second look is almost always failing here.
Jewellery. A tilhari that is half-tilhari half-mangalsutra. A pote that wraps strangely or fades into the neck. Earrings that match on one side and not the other. The model has not learned these as objects; it has learned them as visual texture, and the texture is correct only on average.
Symmetry of faces. Real faces are slightly asymmetric. AI faces are sometimes too symmetric, which produces the uncanny effect a reviewer cannot articulate but the audience absorbs. If a face feels off and you cannot say why, this is usually it.
An extra person or limb in the background. Crowds at festivals are where this hides. The model produces a clear foreground and a smeared background that, on a third look, has a person with two left arms or a child whose legs come out of an adult’s torso. Zoom in on every background figure before approval.
The professional habit is a checklist. Before any AI-generated image with people in it leaves the studio, a reviewer looks at: hands, eyes, jewellery, symmetry, background figures, and the colour palette. Five minutes, a small set of common failure modes, the difference between a campaign and an embarrassment screenshot circulating on Twitter the day after launch.
When to use real photography instead
There is a clean rule for when the model is the wrong tool. If the audience will read the image as “this is a real customer of ours” or “this is a real member of our team,” AI-generated faces are the wrong choice. Full stop. The brand is making a representational claim with that image — this person uses our product, this person works here — and an AI-generated face presented as real is the reputational disaster waiting to happen. When it is discovered, and it will be discovered, the brand becomes the brand that faked its customers.
The rule has clean implications. Hero images that say “meet our customer Maya from Lalitpur” — real photography, with consent, full stop. Team pages — real photography. Founder portraits — real photography. Anything implying a real interaction between a real person and your brand — real photography. The cost of a one-day shoot with a local photographer is small compared to the cost of being the brand that invented its customers.
AI-generated imagery is the right tool where the audience does not assume a representational claim. Mood imagery for the brand. Illustrative backgrounds for a blog post. Conceptual visuals for a thought-leadership piece. Stylised festival imagery for a greeting card that is clearly not a photograph. Anything where the visual register is openly creative rather than documentary.
Concrete prompt patterns
Three Nepal-specific prompt patterns that produce usable output in 2026, given as starting templates you can adapt.
A Dashain greeting card. “A multi-generational family scene during Dashain in a Kathmandu Valley home, the grandmother sitting in a carved wooden seat applying red tika and yellow jamara to the forehead of a young grandchild, four other family members standing around in warm soft natural light from a courtyard window. Older men in daura suruwal with dhaka topi, older women in gunyo cholo and traditional jewellery, younger adults in modern-cut kurtha-suruwal, children in clean simple clothes. Brick courtyard floor, carved wooden windows, marigold garlands across the doorway, a brass household shrine in the corner with oil lamps lit. Warm earth tones, deep reds, brass, marigold orange. Photographic, shallow depth of field, candid moment, no posed expression, no on-screen text. Avoid: Mumbai-style saree, beige kurta on the men, generic Diwali furniture, Holi-coloured powder, neon.”
A tourism campaign image. “A young couple in their late twenties trekking on a stone-paved trail in the Annapurna foothills in mid-morning light, low cloud below the path, terraced fields visible in the middle distance, a small Gurung settlement on the ridge above. The couple in modern technical trekking gear — fleece, light packs — looking out at the view rather than at the camera. The peaks in the far distance are Annapurna South and Hiunchuli, identifiable by shape. Photographic, wide composition, slightly faded colours, photojournalistic style. Avoid: generic Himalayan stock peaks, mountaineering expedition aesthetic, monastery B-roll, prayer flags as decoration.”
A fintech illustration, stylised, no photoreal people. “A flat vector illustration in muted warm tones — terracotta, brass, sage green, soft cream — showing a stylised hand holding a smartphone with a payment confirmation screen visible. The background is an abstract pattern hinting at a Patan courtyard — geometric carved-window motif, suggestion of marigold strings — without rendering any specific human face. Illustration style, no photorealism, no human faces, no text inside the image. Avoid: photoreal phones, stock-app screenshots, anglicised business aesthetics.”
The pattern across all three: heavy specification, named avoid clauses, a clear style register (photographic for the first two, illustration for the third), and a deliberate decision about whether faces are in the frame at all. The fintech illustration deliberately avoids faces because the audience would read photoreal faces in a payment ad as customers — and the brand is not ready to make that claim.
Check your understanding
Quick check
—A marketer types “a happy Nepali family at Dashain” into a current image model. What is most likely to come back, and why?
What comes next
Briefing the model and generating the visuals gets a campaign to the door. Shipping it well is a separate discipline — and one of the most common ways teams misuse AI here is by producing fifty thin variants of an ad instead of two or four that actually test something. The next section is about variation and A/B testing: how to use AI to widen the test space without drowning the signal, and how to read results honestly when traffic is small.