
Southeast Asia’s artificial intelligence adoption is not following the neat, text-first path seen in many early AI markets. Instead, users across the region are jumping straight into a more instinctive mode of interaction: talking to AI, showing it photos, and asking it to work with images, video, and sound. This is the central finding of the “Gemini Report Southeast Asia 2026,” which indicates users in the region now rely on voice, photos, and videos as inputs 42 per cent of the time when engaging with the assistant.
In practical terms, this means AI is becoming less like a search box and more like a companion that can see and hear the world around its users. The shift reflects the region’s internet economy, which has always been unusually visual and mobile-first. Commerce happens on chat apps, and brands are built on platforms like Instagram and Shopee Live. For small sellers who often operate without design teams or copywriters, multimodal AI is not just a productivity tool. It is a shortcut from idea to content, and from content to income.
Indonesia’s Rapid Creative Shift
Indonesia stands out as the region’s clearest example of this behavior. The report describes the country as Southeast Asia’s “Creative Champion,” with 32 per cent of all user journeys classified as creative. Indonesian users are generating around 9 million images every day, the highest volume in the region. This fits neatly into what the report calls Indonesia’s “sat-set” culture, a local expression that broadly means quick, easy, and no fuss.
Related: Singapore slow to adopt cyber threat checks
The behavior captures the way many Indonesian creators and small business owners already work: fast, mobile, practical, and highly responsive to trends. For a micro-influencer in Surabaya, the workflow can be as simple as uploading a photo of cafe decor and asking Gemini to produce a caption within seconds. The same creator might use image generation to build quick mood boards for an upcoming shoot, replacing hours of manual searching and editing with a few prompts.
One example in the report involves an Indonesian handyman in Sidoarjo who uses voice prompts through Gemini Live while repairing washing machines. With greasy hands and limited space, typing is impractical. By asking for error code descriptions aloud while standing in a cramped laundry room, he saves time on each visit and can fit more jobs into a day.
This dynamic suggests that the barrier to entry for digital participation is falling. When a street vendor can produce marketing materials or a repair worker can access technical diagrams without stopping their work, the gap between formal training and practical execution narrows significantly. The technology is effectively outsourcing specialized skills to anyone with a smartphone and a specific problem.
Related: Keys to Success and How to Start a Small Capital Business for Beginners
Older Users in Thailand Drive Adoption
While AI adoption is often framed around Gen Z, Thailand offers a different pattern. The report says users over the age of 54 are the most multimodal age group in the country. They account for 10 per cent of Thailand’s user base and are more likely to engage through voice or images than through typing.
This is a reminder that multimodal design can be an accessibility feature, not just a creative one. For older users, typing long questions on a phone can be cumbersome. Taking a picture or speaking aloud is often easier. One example in the report is a retired teacher in Chiang Rai who photographs local snacks such as Thong Yip at the market to check their sugar content. The tool helps him manage his health more independently between monthly doctor visits.
For Thailand, where the population is aging faster than many of its neighbors, such use cases could become increasingly relevant. AI products designed around voice, images, and local language support may be better suited to older users than text-heavy interfaces built for office workers. At the other end of the age spectrum, the report describes an engineering student in Bangkok using Gemini on her tablet as a second brain. She talks aloud to the AI while riding the BTS Skytrain, asking it to explain complex textbook slides in Thai during late-night study sessions.
Related: Top 4 Ways To Measure Your Home’s Internet Quality Today.
Design Tools and Market Competition
The next phase of this shift could involve deeper integration with design tools. The report highlights a coming integration with Canva, one of the most widely used design platforms among Southeast Asian startups and small businesses. The collaboration would allow users to edit Gemini-generated images and create fuller designs directly within the AI interface.
For founders and solopreneurs, that could reduce a familiar bottleneck. A small business owner could move from a rough campaign idea to a usable flyer or social post without switching between multiple apps. Google, however, is not alone in chasing this market. OpenAI’s ChatGPT has pushed aggressively into multimodal interaction, including image understanding and generation. Meta is embedding AI features across Facebook, Instagram, and WhatsApp, platforms that remain central to Southeast Asian commerce and content distribution.
TikTok owner ByteDance also has a strong position through its creator tools and video-first ecosystem. In this market, distribution may matter as much as model quality: the winning tools will be those already sitting inside the apps people use every day. For the region’s startup ecosystem, the lesson is clear. AI adoption in Southeast Asia will not be won through English-only chatbots designed for desktop workflows. The stronger opportunity lies in tools that understand local languages, work well on mobile, and fit into the rhythms of everyday work.
