First-party data
What Language Malaysian Facebook Ads Use (2026)
The language split across 221,588 running Malaysian Facebook and Instagram ads, merged from 374 raw labels, cut by vertical, with the full merge map and dataset as a free CSV.
Updated October 2026 · AdPlay.ai Team
Malaysian Facebook and Instagram ads run mostly in English, but not by as much as you would guess. Across 221,588 running Malaysian ads carrying a language label on 25 August 2026, 45.6% are in English, 27.1% in Malay and 22.7% in Chinese, with Tamil almost absent at 56 ads. The mix moves sharply by vertical: Malay carries 46.7% of transport ads against 16.4% in fitness. One label per ad means code-switching is invisible here, so every multilingual figure on this page is an understatement rather than an estimate.
Everyone advertising in Malaysia has an opinion about which language to run in and nobody has published the split. This page counts it: every distinct language label in the archive, merged into families by a map you can read and disagree with, across the ads actually running. It also tests one of our own claims, because a page that only checks other people's numbers is not doing the harder half of the job.
English leads, and the other half is not one language
45.6% of running Malaysian ads carry an English label, 27.1% Malay and 22.7% Chinese. The first number is lower than most people expect and the third is higher. The practical consequence is that English is a plurality rather than a default. More than half of Malaysian advertising is not in English, and it splits into two audiences that need different creative rather than a translation. An advertiser who runs English only is not covering the market with a lingua franca, they are competing in the most crowded of three lanes. Two smaller numbers are worth sitting with. Tamil accounts for 56 ads out of more than two hundred thousand, which is close to absent for a language spoken by a substantial community. And 1.9% of ads carry a language that is not a market language here at all, mostly Spanish, Turkish, Korean and Thai. Those are almost certainly Malaysian-registered advertisers selling into other countries, which is a real and separate observation rather than a data error, and it is why those labels are mapped deliberately rather than dropped.
| Language family | Ads | Share of ads carrying a label |
|---|---|---|
| English | 101,139 | 45.6% |
| Malay | 60,086 | 27.1% |
| Chinese | 50,288 | 22.7% |
| Other | 4,244 | 1.9% |
| Multiple | 3,753 | 1.7% |
| Indonesian | 1,489 | 0.7% |
| Arabic | 523 | 0.2% |
| Tamil | 56 | 0% |
The vertical is a better predictor than the country
The national split hides most of what is useful. Transport ads run 46.7% Malay against 36.4% English. Fitness inverts it at 16.4% Malay and 57.7% English. Chinese moves less at the extremes but sits above a quarter of ads in several categories. Read that as audience rather than as language preference. Categories that sell into a broad national market skew Malay; categories that sell to an urban, higher-income or younger buyer skew English; categories with a strong Chinese-Malaysian customer base carry a Chinese share far above the national average. None of that is surprising to anyone who sells here, and until now none of it had a number attached. What this does not license is a rule like "use Malay for services and English for software". The figures describe what advertisers currently do, and the archive holds no spend or results, so a category running mostly in one language is evidence about convention rather than about what converts.
| Vertical | English | Malay | Chinese |
|---|---|---|---|
| Transport | 36.4% | 46.7% | 12.2% |
| Finance | 49.8% | 37.9% | 8.9% |
| Services | 35.7% | 36% | 24.3% |
| App / Software | 42% | 30.4% | 17.3% |
| Travel | 56.8% | 27.9% | 11.5% |
| Healthcare | 31% | 27% | 38.8% |
| Retail / Marketplace | 47.6% | 26.7% | 19.5% |
| Real Estate | 55.4% | 24.6% | 17% |
| Home / Garden | 50.3% | 24.5% | 21.6% |
| Clothing / Apparel | 58.9% | 24.4% | 7.5% |
| Beauty / Personal Care | 40.1% | 23.8% | 31.3% |
| Electronics / Tech | 58.1% | 23.2% | 14.1% |
| Food / Beverage | 35.2% | 22.4% | 38.1% |
| Consumer Goods | 48.9% | 21.9% | 23.7% |
| Education | 44.8% | 20.9% | 27.8% |
| Kids | 47.1% | 20.9% | 25% |
| Fitness | 57.7% | 16.4% | 21.4% |
The merge is a judgement, so here it is
The archive stores language as free text written by a model, and it holds 374 distinct values for what is really a handful of languages: Chinese arrives as Chinese, as zh, as zh-CN, as Traditional Chinese, as Chinese (Simplified) and as Mandarin. Producing a split at all means merging them, and merging is a judgement. So the judgement is published. The merge map is a file in the repository, every raw label ships in the CSV as its own row with the family it was folded into and the rule that did it, and the builder refuses to run if a label above a size threshold is not in the map deliberately. That last part is not decoration: it fired on the first run and named fourteen languages that were about to fall through silently. One rule worth stating because it is the easy mistake: a label naming two languages, like English/Malay or Mixed (English and Chinese), is counted as Multiple and never split between families. There is no way to know from a single label which language the ad led with, so splitting it would be inventing the answer.
| Raw label | Merged into | How it was merged | Ads |
|---|---|---|---|
| English | English | Named in the map | 99,513 |
| Malay | Malay | Named in the map | 59,879 |
| Chinese | Chinese | Named in the map | 45,509 |
| zh | Chinese | Named in the map | 3,401 |
| en | English | Named in the map | 1,626 |
| Spanish | Other | Named in the map | 1,093 |
| Traditional Chinese | Chinese | Named in the map | 1,025 |
| Indonesian | Indonesian | Named in the map | 992 |
| English, Chinese | Multiple | Names two or more languages | 508 |
| id | Indonesian | Named in the map | 497 |
| Arabic | Arabic | Named in the map | 494 |
| Malay/English | Multiple | Names two or more languages | 390 |
| English/Malay | Multiple | Names two or more languages | 385 |
| French | Other | Named in the map | 380 |
| Korean | Other | Named in the map | 348 |
Proving the list is complete, and testing one of our own claims
A language study is only as good as its list of languages, and getting that list is harder than it sounds. The search index returns at most 100 distinct values for a field, alphabetically, so one query gives a list that looks complete with an arbitrary tail missing. Partitioning by vertical, then by format, then by year brings 73 partitions back under the cap, and 4 still truncate at the deepest split used. Rather than claim completeness we measured the gap. Counting every enumerated label exactly and subtracting the sum from the ads known to carry any label leaves a residual of 10 ads out of 221,588. The list is complete to within ten ads, and that residual is a published row rather than a footnote. The study also had a target closer to home. Our own skincare hub asserts, with nothing behind it, that in Malaysian skincare "English leads, Chinese is a close second, and Malay carries a meaningful share". Measured across beauty and personal care ads: English 40.1%, Chinese 31.3%, Malay 23.8%. The claim holds, including the ordering, which is a better outcome than we had any right to expect from a sentence written on instinct.
What this cannot tell you
The limit that matters most is one label per ad. Malaysian ad copy code-switches constantly, and a caption that opens in Malay and closes with an English call to action carries a single label. So the Multiple family at 1.7% is a floor on multilingual creative and nothing like a measurement of it, and the single-language shares are correspondingly inflated. Nothing in the data can fix this; only re-reading the copy itself could. The label is also model-written rather than declared by the advertiser, so it is an inference from the creative. Read a gap of twenty points as real and a gap of two as noise. And one thing this page is explicitly not about: it describes ad creative, not websites. It says nothing about which language a Malaysian business should publish its site in, and it is not an argument for or against a Malay-language version of anything.
Reuse this data
Every raw label, every family total, the per-vertical cut and the coverage residual are in one CSV under CC BY 4.0, free to reuse including commercially with attribution to AdPlay.ai. Because the merge map ships inside the file, you can rebuild the families a different way from the same rows if you disagree with ours, which is the point of publishing a judgement rather than a conclusion.
By the numbers
Frequently asked questions
What language are most Malaysian Facebook ads in?
English, but only just: 45.6% of running Malaysian ads carry an English label, against 27.1% Malay and 22.7% Chinese. More than half of Malaysian advertising is not in English.
How many Malaysian ads run in Chinese?
22.7%, or 50,288 of the 221,588 running ads carrying a language label. It is higher than most advertisers assume and rises well above a quarter of ads in several verticals.
Should I run my Malaysian ads in Malay or English?
This data cannot answer that, because the archive holds no spend, clicks or results. What it can tell you is what your category currently does, which is a guide to how crowded each lane is rather than to which one converts. The per-vertical table is the useful part.
Why is Tamil almost absent?
Tamil accounts for 56 ads in the archive. That is a real observation about advertising rather than about the population, and it is the kind of gap worth noticing if you sell to a community nobody is speaking to in their own language. Bear in mind the one-label-per-ad limit: a code-switched ad carrying a Tamil line would not be labelled Tamil.
How can one ad have two languages?
It can, and that is this study's biggest blind spot. The field holds one label per ad, so an ad that opens in Malay and closes in English is recorded once. Labels naming two languages are counted as Multiple, at 1.7%, and that figure is a floor on code-switching rather than a measurement of it.
Why do some Malaysian ads run in Spanish or Turkish?
1.9% of ads carry a language that is not a market language in Malaysia, mostly Spanish, Turkish, Korean and Thai. The most likely explanation is Malaysian-registered advertisers selling into other countries. Those labels are mapped deliberately rather than discarded so the figure stays visible.
How do you know you found every language label?
We do not claim to, we measured the gap instead. The index returns at most 100 distinct values per query in alphabetical order, so the labels were enumerated by partitioning until partitions came back under the cap, then counted exactly. The sum falls 10 ads short of the ads known to carry a label, and that residual is published as its own row.
Can I reuse this dataset?
Yes, under CC BY 4.0 with attribution to AdPlay.ai. The CSV carries every raw label alongside the family it was merged into and the rule that did it, so you can redo the merge your own way from the same rows.
The data behind this page
Every figure on this page is a row in that file. Measured as at 2026-08-25. Published under Creative Commons Attribution 4.0, so reuse it with a credit to AdPlay.ai.
Sources
Keep exploring
Turn ad research into winning ads
See what 16,000 Malaysian brands advertise, then generate on-brand creative, all in one tool.
7-day free trial · No credit card required
