First-party data

What Our Malaysian Ad Archive Can and Cannot Tell You

The corpus card for the AdPlay.ai Facebook ad archive: how many Malaysian ads it holds, what window they cover, how completely each field is populated, and the list of questions it can never answer, with the whole audit as a free CSV.

Updated September 2026 · AdPlay.ai Team

Quick answer

The AdPlay.ai archive holds 1,043,499 Malaysian Facebook and Instagram ads out of 1,428,677 tracked across all markets, 222,921 of them still running as at their own last check on 25 August 2026. It is a tracked set, not a census: collection follows advertisers rather than a sampling frame, so every count is a floor. 98% of Malaysian ads in it started between 2025-05-25 and 2026-08-21, and the median started 2026-04-20. It carries creative metadata only, so it cannot tell you what any ad cost, who saw it, or whether it worked, and this page publishes the full list of what it cannot answer alongside what it can.

Every other study on this site quotes numbers out of one archive. This is the page that says what that archive actually is. It is deliberately the least exciting page here: counts, a date range, a field-by-field coverage audit, two contamination checks, and a list of questions we refuse to answer from this data. If you are deciding whether to cite anything else we publish, read this first, and if a figure elsewhere on the site contradicts this page, this page is the one to trust.

What is in it, and over what window

The archive holds 1,428,677 Facebook and Instagram ads across all markets, of which 1,043,499 carry an advertiser classified as Malaysian. 222,921 of those were still running the last time we checked that specific ad. The start-date window matters more than the totals do. 98% of the Malaysian ads sit between 2025-05-25 and 2026-08-21, with a median start of 2026-04-20, so this is a recent archive rather than a historical one. That is a real limit on what can be asked of it: a question about how Malaysian advertising changed over several years cannot be answered here, and any figure framed as a trend across that span would be reading collection growth as market change. The deeper limit is that collection follows advertisers we track rather than a sampling frame. There is no version of this dataset that is a census of Malaysian advertising, and no weighting that would turn it into one, so every count published anywhere on this site is a floor. Where a page says how many ads do something, the honest reading is always at least this many.

What the AdPlay.ai Malaysian ad archive holds, 2026
Document counts and ad start-date coverage for the Malaysian slice of the AdPlay.ai Facebook ad archive, as at 25 August 2026.
MeasureValue
Ads tracked across all markets1,428,677
Malaysian ads1,043,499
Malaysian ads running at their last check222,921
Earliest start date, first percentile2025-05-25
Median start date2026-04-20
Latest start date, ninety-ninth percentile2026-08-21

Which fields are actually populated

A field that exists on the schema is not a field you can count with. The sample below is 2,000 Malaysian ads drawn evenly across the whole start-date range rather than off the recent end, because sampling the newest ads would measure the latest collection pass instead of the archive. Three results change what may be asked. Ad body copy is effectively universal at 99.8%, so a keyword study over body text has an honest denominator. The English renderings do not: bodyEnglish sits at 34% and hookEnglish at 33.4%, because they exist only where the original was not already English. A numerator drawn from those fields against a whole-corpus denominator is not a floor, it is uninterpretable, and that rules out a whole class of question. Video transcription is present on 36.9% for the same structural reason, and it was removed from the search index besides, so no keyword count can reach it at all. The model-generated labels are the opposite case: creative angle and industry are attached to 100% of sampled ads. Universal coverage is not the same as accuracy, though. They are model output, they are multi-label, and the classifier has changed over the archive's life, so they carry cross-sections and never trends.

Field coverage in the AdPlay.ai Malaysian ad archive, 2026
Share of sampled Malaysian ads carrying a non-empty value in each field, over 2,000 ads drawn evenly across the archive's whole start-date range, as at 25 August 2026.
FieldAds carrying a value
body99.8%
title76.3%
visualText60.3%
hook99.9%
bodyEnglish34%
hookEnglish33.4%
transcription36.9%
language98.7%
publisherPlatforms100%
ctaType94%
imageUrl57%
videoUrl43.8%
theme100%
industry100%

Three fields you cannot list in one query

The search index returns at most 100 distinct values for any field, and it returns them in alphabetical order rather than by frequency. On a field with more values than that, one query gives you a list that looks complete and has an arbitrary tail missing, with nothing in the response saying so. Measured here, theme, language, brandName all exceed the cap. That has direct consequences: the advertiser field cannot be enumerated in one pass, so a question about how many distinct Malaysian brands advertise needs a partitioned crawl rather than a single query; creative angle has to be counted against the classifier's own controlled vocabulary rather than by asking what values exist; and any language study has to partition until every partition comes back under the cap and then prove the parts add up to the whole. Industry, ad format and publisher platform stay under the cap and can be listed directly. Every study we publish either respects that or does not ship. The tooling behind these pages now refuses a facet response that comes back at exactly the cap, because a list that is silently short is worse than an error.

Two checks on the start date, one of which mattered

Run length and every seasonal timing question depend on one field, so it gets checked rather than trusted. The code that writes it falls back to the current time when Meta supplies no start date, which would stamp an ad with the moment we ingested it. Fake starts of that kind would cluster on the days we happened to be collecting, and in a study that buckets ad starts by week they would read as a launch ramp that never happened. It cannot be detected by filtering, because the ingest timestamp is not a queryable field. It can be detected by arithmetic. A real start date is a calendar day rendered at a fixed offset, so its remainder against a full day is the same constant for every ad written that way, while an ingest stamp carries whatever wall-clock time it happened to be written at and is therefore very nearly unique. Across the sample, 2 offsets account for every ad, 7 hours on 1,526 and 8 hours on 474, and the share carrying an offset that almost nothing else shares is 0%. The fallback is effectively never exercised on Malaysian ads, so weekly bucketing of start dates is sound. The second check is less comfortable. Ad format is written by two different code paths that disagree about whether a dynamic creative is filed as dynamic or as an image, so the boundary between those two values is writer-dependent. Any split between video and static on this data is a range rather than a point, and every page that cuts by format says so.

What it cannot answer, and why

The list below is the part of this page worth citing. Most published ad datasets describe what they contain; almost none publish what they structurally cannot measure, which is what a reader actually needs in order to know whether a number is safe to repeat. The short version: no spend, no cost per anything, no impressions, no reach, no return, ever. Those fields are not in Meta's public Ad Library, so any figure of that shape derived from this archive would be invented, and any Malaysian benchmark quoting them without naming its source deserves the same suspicion. No targeting, because it is not published and the audience fields here are model inferences from the creative. No claim that an ad performed, because run length is right-censored and says only that an advertiser kept paying. Every one of those refusals is a row in the CSV with an empty value and its reason attached, so a machine reading this dataset gets the limits along with the numbers.

Questions the AdPlay.ai ad archive cannot answer, 2026
Questions this dataset is structurally unable to answer, why each one is out of reach, and what we publish in its place, as at 25 August 2026.
QuestionWhy it cannot be answeredWhat we publish instead
What does a Facebook ad cost in MalaysiaMeta publishes no spend, impressions or reach in the public Ad Library, so no CPM, CPC, CPA or budget figure can be derived from anything hereDated third-party cost sources, cited and linked on the guides that quote them
Which ads performed bestNothing in the archive measures outcomes. Run duration is the only signal, and it is right-censored, so it says an advertiser kept paying rather than that an ad workedHow long ads run, and what the ones that last have in common
How much of the Malaysian market this coversCollection follows advertisers we track rather than a sampling frame, so the archive is a tracked set and not a census. Every count is a floorThe absolute counts, always with the snapshot date beside them
Which ads were shown to whomTargeting is not published. The audience fields are model inferences from the creative, not delivery dataNothing. This is not estimated
Whether an ad is running right nowLiveness is as at each ad own last check, and the refresh cadence is uneven, so an ad last checked in April is counted as it stood in AprilThe confirmed-to date beside every run length
How advertisers split budget across placementsThe publisher-platform field records eligibility, which is what an advertiser allowed, not where delivery or spend wentEligibility shares, labelled as eligibility
Which language an ad was written in when it mixes twoLanguage is one free-text label per ad, so code-switching is invisible and every multilingual share is understatedThe raw label distribution with the merge map published beside it
How a category performed over timeIndustry and creative-angle labels are model-assigned and multi-label, and the classifier has changed over the archive lifetime, so a time series across them would measure the modelPoint-in-time cross-sections, each dated

Reuse this data

The whole audit is one CSV under CC BY 4.0: counts, the coverage window, per-field population rates with the sample size attached, the enumeration limits, both contamination checks, and the refusal rows. Use it including commercially with attribution to AdPlay.ai. If you cite a figure from any of our other data pages, cite this one beside it, because the caveats live here rather than being repeated in full on every page.

By the numbers

1,043,499
Malaysian ads in the archive
AdPlay.ai archive, 25 August 2026
1,428,677
ads tracked across all markets
AdPlay.ai archive, 25 August 2026
99.8%
of sampled ads carry ad body copy
AdPlay.ai archive, 25 August 2026
34%
of sampled ads carry an English rendering of their body copy
AdPlay.ai archive, 25 August 2026
98%
of Malaysian ad start dates fall inside the published window
AdPlay.ai archive, 25 August 2026
0%
of sampled start dates look like an ingest timestamp rather than a real start date
AdPlay.ai archive, 25 August 2026

Frequently asked questions

How many Malaysian Facebook ads does the AdPlay.ai archive hold?

1,043,499 as at 25 August 2026, out of 1,428,677 tracked across all markets, with 222,921 still running as at each ad's own last check. It is a tracked set rather than a census, so treat that as a floor rather than as the size of Malaysian advertising.

What period do the ads cover?

98% of the Malaysian ads started between 2025-05-25 and 2026-08-21, and the median start date is 2026-04-20. It is a recent archive, so it supports point-in-time cross-sections and does not support multi-year trends.

Can this data tell me what Facebook ads cost in Malaysia?

No, and no dataset built on the public Ad Library can. Meta does not publish spend, impressions or reach there, so cost per thousand, cost per click and return on ad spend cannot be derived from it. Any Malaysian cost benchmark you find should say where its numbers came from, and several widely republished ones do not.

Is a long-running ad a successful ad?

It is evidence, not proof. An ad is recorded as running because it has not stopped yet, not because it completed a successful run, so run length is right-censored. What it tells you is that an advertiser chose to keep paying, which is meaningful because ad budgets get cut quickly, and which is not the same as a measured result.

Why can you not simply list every advertiser in the archive?

Because the search index returns at most 100 distinct values for a field, in alphabetical order rather than by frequency, so one query on the advertiser field gives a list that looks complete with an arbitrary tail missing. Counting distinct advertisers honestly needs a partitioned crawl that proves the parts add up to the whole, which is a different piece of work from a single query.

How do you know the ad start dates are real?

We checked rather than assumed. The writing code falls back to the current time when Meta gives no start date, so we measured how many start dates look like a timestamp of when we collected the ad rather than a calendar date. Across 2,000 sampled ads the answer is 0%, because a real start date renders at one of a small number of fixed offsets while an ingest timestamp is very nearly unique.

Are the industry and creative angle labels reliable?

They are model-generated, attached to 100% of sampled ads, and multi-label, so an ad can carry several and shares across them exceed 100%. Read a large gap between two labels as real and a small one as noise, and do not read a change over time as a market change, because the classifier itself has changed over the archive's life.

Can I reuse this audit?

Yes, under CC BY 4.0 with attribution to AdPlay.ai. The CSV carries every count, the coverage window, the per-field population rates with their sample size, the enumeration limits and the refusal rows, so you can check any figure on this page or cite the limits alongside a number you take from elsewhere on the site.

The data behind this page

Download the full dataset (CSV)

Every figure on this page is a row in that file. Measured as at 2026-08-25. Published under Creative Commons Attribution 4.0, so reuse it with a credit to AdPlay.ai.

Sources

Keep exploring

Turn ad research into winning ads

See what 16,000 Malaysian brands advertise, then generate on-brand creative, all in one tool.

7-day free trial · No credit card required