Every influencer marketing team eventually hits the same wall, and it is not the one they planned for. The budget is approved, the brief is written, the reporting dashboard is ready. Then somebody asks a simple question: who are we actually working with?
That question turns out to be the most expensive part of the whole program. In the 2025 industry benchmark report, finding the right creators was named the single biggest in-house challenge by 30 percent of brands, ahead of measuring ROI, managing contracts, and processing payments combined. A separate market review of the creator economy put the same figure higher, with 48 percent of marketers calling creator discovery their biggest challenge, and market statistics published by the same research group show discovery and vetting is the function most often handed to an outside agency.
So the money flows, and the matching does not. Global influencer marketing spending passed 33 billion dollars, with the tools segment alone heading toward two billion by 2031 on one estimate and a steeper path on another. Yet the search remains the weakest link in the chain. This article is about that gap: where brands actually look, why the search breaks, what vetting really costs, and what a discovery process that scales looks like in practice.
Why the influencer marketing discovery problem is structural
Most marketing decisions are made with data you already own. Your ad account knows who clicked. Your store knows who bought. Discovery is the opposite: you are shopping in a market where what you are buying is a relationship with someone whose audience you cannot fully see, and whose quality you cannot verify from the outside.
Economists have a name for this. It is a matching market with high search costs and unresolved quality uncertainty, and the predictable result is that buyers stop comparing and start defaulting to whatever signal is cheapest to read. Alvin Roth's work on market design describes how matching markets behave when participants cannot evaluate each other cheaply, and platform economics research shows how intermediaries emerge to lower exactly that cost. The same dynamic shows up in digital marketplaces far from marketing: management research on platform strategy and economic analysis of marketplace efficiency both describe the intermediary as a search-cost machine.
In creator marketing, the cheapest signal has always been follower count. It is visible on every profile, it correlates loosely with reach, and it lets a buyer decide in ten seconds. It is also a blunt instrument. Research in PLOS ONE on how consumer trust forms found that audience size predicts campaign outcomes inconsistently at best, and a meta-analysis in the Journal of Business Research reached the same conclusion from a different angle: the signals that travel furthest in a pitch deck are rarely the ones that drive results. Business school research on marketing strategy has made the broader version of this point for years, and studies indexed in PubMed on engagement and disclosure reach it from the audience side.
Reputation systems exist to solve it. A survey of trust and reputation systems in Decision Support Systems laid out the mechanics two decades ago, and the pattern holds in creator marketing: where verifiable evidence is absent, buyers fall back on fame as a proxy for quality. That proxy is what most discovery tooling quietly encodes.
Then came the natural experiment. When a major social platform in Asia launched a centralized marketplace that let brands search, filter, and compare creators directly, researchers Shi, He and Ba tracked more than 152,000 posts from 1,161 creators before and after the launch. Sponsorship spending did not keep concentrating on the biggest names. It shifted toward the middle tier, because the marketplace replaced reputation with something verifiable, and a plain-language summary of the paper walks through the mechanism.
Read that finding carefully, because it is the whole ballgame. Transparency did not make famous creators more findable. It made unknown creators safe to bet on. Search cost, not talent, was the barrier. Consumer behaviour research on influencer marketing strategy reaches a compatible conclusion: the industry's binding constraint is not supply of creators, it is the cost of choosing between them.
What discovery actually involves
Teams tend to talk about discovery as one task. It is really three, and they fail in different ways.
Finding. Generating a candidate pool wide enough that the right creator is in it at all. Fail here and no amount of clever vetting helps, because the good option was never on the list. This is the part artificial intelligence has genuinely changed, and adoption data on AI inside creator programs shows discovery as the most common use case by a wide margin, well ahead of brief writing and reporting.
Qualifying. Deciding who deserves a look, using signals that survive contact with reality: view consistency, comment quality, audience geography, sponsorship history, brand safety. A practitioner on Reddit described building exactly this split into a two-layer method: what you can see publicly, then what you have to infer from behaviour. Research on how source credibility and content quality interact gives the general case for why both layers are needed.
Where brands actually look for creators
Strip away the tooling and there are four search routes in real use. They are not equivalent, and most teams lean on the weakest one first.
Native search on social platforms
Typing topic keywords into a social platform and filtering to people rather than posts. It is free, fast, and reflects what is resonating right now rather than what was indexed months ago. The catch is that social platforms were built for content search, not people search. Filtering by audience demographics, language mix, sponsorship history, or engagement quality is either impossible or unreliable, and the official help documentation is written for creators monetising content, not for buyers selecting partners. You get names, not evidence. The platform's own creator hub and the creator education academy are useful for the supply side and almost silent on the demand side.
Why creator databases only solve half the problem
Creator databases
Subscription tools that index large numbers of profiles with filters for audience, performance, and growth. These solve breadth. The limitation is that they answer who exists, not who is right for this brand, at this budget, on this brief. Two separate questions dressed up as one. Databases also go stale between crawls, and a benchmark study of discovery workflows found that automated agents handle the compile-and-list part well while still letting fraudulent accounts through the final shortlist. Review sites have noticed the same pattern: the creator economy data published this year argues that scrape-based indexes go stale precisely because they never ask creators for their own data.
Competitor sponsor history
The highest-leverage route, and the most underused. If a creator has already promoted a competing product, three things are almost certainly true: their audience overlaps with yours, they have been through a brand workflow, and they can hold a sponsored segment without making it feel like an advertisement. Research on sponsorship disclosure and research on credible attributes and parasocial trust explains why the second point matters, and research on how audiences respond to sponsored recommendations separates a segment viewers tolerate from one they skip.
Reading sponsor history properly means reading descriptions and transcripts, not labels. Self-reported paid promotion markers are unreliable, which is both a compliance problem and a reason databases built on those markers miss real deals. The rules on endorsements and testimonials are unambiguous about disclosure, and legal commentary on the material connection standard makes it clear where the line sits. The practitioner guidance is worth re-reading once a year, and trade coverage of enforcement actions is a useful reminder that the rules are being applied.
The warmest list you already own
Your own data
Customers and community members already creating content about the product. This audience is pre-qualified for affinity, usually cheap to reach, and often skipped entirely because it sits in a different team's data. Research on word of mouth has pointed this way since Sernovitz's work on how smart companies get people talking, and the academic base is deep: work in the Journal of Consumer Research on how word of mouth shapes decisions, research in the Journal of Marketing Research on electronic word of mouth and brand evaluation, and a Journal of Marketing study on turning everyday customers into advocates. Employee advocacy benchmarks apply the same logic inside a company, where the audience is warm by default.
What search results cannot tell you
Here is a concrete test. Take any shortlist you have built in the last month and ask which of these you can answer without leaving the tool.
Can you see who else has sponsored this creator, and how often? Can you see whether their sponsored posts perform near their organic average, or collapse? Can you see the relationship record, meaning which brands they have delivered for and whether those partnerships repeated? Can you see median performance across recent posts rather than a lifetime average distorted by one viral hit?
For most teams the honest answer is no on the first three, and a partial yes on the fourth. This is the structural weakness of filter-based discovery. Filters describe a profile. They do not describe a relationship, and a relationship is what you are buying. Selection research on consumer trust in recommendations and on research on online trust and perceived risk in social commerce keeps landing on the same point.
Median versus average is worth dwelling on, because it changes decisions. A creator with a modest following whose last nine posts each did between three and seven thousand views and whose tenth did four hundred thousand has an average that looks impressive and a median that tells the truth. If you plan against the average, you overpay and underdeliver. If you plan against the median, you know what you are buying. The economics of matching markets, as set out in working papers on platform economics and open access economic research, come down to this: the quality of a match depends on the quality of the evidence you can bring to it. The same logic appears in statistical reporting on digital services from the national accounts, where measurement choices change the story entirely.
The vetting bottleneck: thirty minutes and a feeling
Once a shortlist exists, vetting is the next place the pipeline jams. Survey work on brand safety found that more than half of marketers spend thirty minutes or less vetting each creator. Manual review of social content is part of the process for 81.2 percent of them, and only 9.4 percent outsource it fully. Asked what is hard, they name time (38.5 percent), sustained monitoring (34.2 percent), and the absence of automation (28.2 percent). Barely one in ten describes their own vetting process as very scalable.
Thirty minutes sounds reasonable until you price it. It is roughly one working day to assess forty candidates properly, which is about what you need to fill a shortlist of eight. And thirty minutes covers a fraction of a creator's output, which is why manual reviewers miss bought followers even when they are explicitly looking for them.
The documentation gap makes it worse. Almost every brand wants a record of how creators were vetted before a campaign runs, and only about a quarter always receive one. Reporting on the maturation of the industry describes the shift from vibes-based casting to data-driven casting as one of the clearest signs the channel is growing up, and the wider trade coverage of creator marketing tracks the same progression. Agencies are not hiding anything sinister; the process is simply manual, undocumented, and different every time.
Influencers feel the same friction from the other side. A thread on the most frustrating part of vetting landed on the same conclusion from the creator's perspective: engagement-based screening is slow and unreliable, and the better proxy is which brands keep working with the same person. Buyers make the same argument from the other chair, from shortlisting creators on evidence rather than guesswork to complaints about the hours spent manually checking follower quality.
The supply side of that market is bigger than most brand teams assume, which makes the screening problem harder rather than easier. Practitioner estimates put the number of people making a living from creator work in the tens of millions, while earnings analysis in a 2026 creator economy study found that nearly half earn under ten thousand dollars a year. A market that large, with that much variance in quality and professionalism, cannot be screened by hand.
The fraud tax on discovery
Discovery and authenticity are the same problem wearing different clothes. You cannot fix matching without also fixing verification, because the signal you are matching against may be manufactured.
The numbers are not small. An analysis of 100,000 accounts found that 37.2 percent of followers showed signs of being fake, purchased, or inauthentic, and roughly 19 percent of total spend reached audiences that were not real. The same body of research puts annual influencer fraud losses in the billions, with fake or bot followers accounting for the majority of reported incidents and enforcement activity rising sharply.
Detection is where the industry is furthest behind. Adoption data shows discovery as the most automated step and fraud detection as the least adopted, with a large majority of marketers rejecting synthetic creator substitutes outright while continuing to under-invest in verification of the human ones. Teams have automated the cheap part of the funnel and left the judgment-heavy part manual. Discovery at scale without verification at scale is exactly how inflated accounts get through, which is the same failure mode technology reporting has documented across platform integrity, platform policy, and systems research.
Academic work on fake follower detection has been available for years, including the fame for sale line of research on detecting purchased audiences and network analysis research on identifying coordinated inauthentic behaviour. The tools exist. The habit does not. Mainstream coverage of these problems, from the Guardian and the BBC to NPR and CNBC, has been steady for a decade, and academic commentary keeps pointing at the same structural weakness.
Why authenticity is a pricing lever, not a nice-to-have
There is a commercial reason to care about verification beyond wasted spend. Consumer research from 2026 found that 85 percent of consumers would pay more for brands they consider authentic, and 93 percent said authentic engagement is what builds trust. The same study found that peer voices and independently verifiable content outrank brand-produced messaging by a wide margin, a finding confirmed in the release summary, broken down by content type, and published in full for teams that want the methodology.
Put those findings together and the discovery problem stops being an operational annoyance. Choosing an inauthentic creator does not just waste budget: it puts your brand inside a context consumers are actively screening for. Vetting is where that risk is managed, and it is the step with the least tooling, the least documentation, and the least time allocated. Consumer psychology research on persuasion and credibility, on why things spread, and on trust in online sources has made that point for decades. Psychological research on social influence and practitioner reporting on persuasion keep confirming it, and peer reviewed work across the behavioural sciences supplies the mechanism.
Quick quiz: how good is your creator discovery process?
Answer honestly, then check yourself.
1. Your team spends under thirty minutes vetting each creator. What is the most likely consequence?
- A. You save money and move faster
- B. You miss creators who look weak on average views
- C. You screen only a fraction of each creator's output, so inflated accounts pass
Reveal the answer
C. Thirty minutes covers a small slice of a creator's content history, which is why manual review misses fake engagement even when the reviewer is deliberately looking for it. Structured verification, with the evidence attached to the creator's profile, is what closes that gap.
2. A creator's average views are strong but their last eight posts were flat. What does that tell you?
- A. Nothing, averages are what matter
- B. One viral post is distorting the average, so plan against the median
- C. You should offer a higher rate
Reveal the answer
B. Median recent performance is the number you can plan against. It also separates the creators who are genuinely worth a premium from the ones a single spike made look bigger than they are.
3. Which discovery route is most likely to surface a creator who can carry a sponsored segment well?
- A. Keyword search on a social platform
- B. A database sorted by follower range
- C. Sponsor history and past brand relationships
Reveal the answer
C. Prior brand work is the strongest available predictor of commercial readiness, because it shows the creator has delivered inside a sponsorship workflow before. That history only helps if the relationship record is visible to both sides.
What an agency actually does differently
Watch a competent agency run discovery and you notice the sequence changes. They start from the audience, not the category. Instead of we need creators in our space, they start with the buyer problem and reverse into the searches that buyer would run: comparison content, tutorials, routines, setup tours, mistake videos. Those formats carry brand messages naturally. Pure entertainment rarely does, and research on parasocial relationships with creators and on messenger and audience congruence shows why the fit between voice and product matters more than raw reach. Research on influence in professional buying extends the same principle to business audiences, where the tolerance for a poorly matched voice is lower.
Then they cast wide on purpose. A longlist of around a hundred candidates, filtered to twenty on evidence, deep-evaluated, and pitched to fifteen or so, is a realistic funnel if the targeting was sound. Teams that pitch two hundred creators with a template do worse than teams that pitch fifteen with real research behind each one. Practitioners who have sat on the brand side of the inbox work on the value chains behind creator deals describes the same pattern: the proposals that get answered are the ones that make evaluation easy.
And they treat the shortlist as an asset, not a spreadsheet. The advice from experienced buyers in creator communities keeps coming back to the same point: ignore follower count, look at comment quality. Creator-to-creator referrals shortcut much of the rest, because creators know whose audiences are real and whose numbers are inflated. Good negotiation discipline helps too, and the classic literature on leverage and information asymmetry in Machiavelli's Prince, Sun Tzu's Art of War, and Adam Smith on how markets price information still reads surprisingly current in a rate negotiation.
The relationship record nobody keeps
One pattern shows up in every serious audit of creator programs: most brands still run one-off activations even though a large share say they prefer long-term partnerships, and repeated collaborations tend to outperform single posts. Every one-off deal throws away the most valuable output of the process, which is knowing how that creator performed for you. Research on the future of social media in marketing, on online influencer marketing as a discipline, and on the research agenda for digital marketing all point at relationship depth as the under-measured variable.
This is the part of discovery that generic search cannot solve, because it belongs to you. A platform that only helps you find people leaves you to remember who worked, who missed deadlines, who delivered early, whose sponsored post outperformed their organic baseline. Most teams remember none of it, so the next campaign restarts from zero. Trade reporting has been documenting that restart cost for years, whether you read advertising trade coverage, marketing industry analysis, agency side commentary, martech reporting, or practitioner resources.
Platforms built around the full deal lifecycle treat this differently. Infmap, for example, gives every user a public profile with performance data and a personalised CRM that keeps deal history attached to the relationship itself, so the second collaboration starts with evidence instead of memory. Discovery, negotiation, contract, and delivery sit in one workflow rather than four disconnected tools, and the payout runs through the same system the deal was agreed in. That architecture matters little for a single campaign and a great deal once a team is running dozens, which is also where customer data platform research and enterprise commerce insights predict the same consolidation.
Is AI fixing discovery or just speeding up the wrong part?
Artificial intelligence is now the default first pass in creator discovery. That is a rational place to start, because sifting thousands of accounts into a shortlist is the most repetitive task in the workflow and the cost of a wrong first pass is low when a human still picks the final list. Reported adoption of AI inside influencer marketing programs is now high, with a small minority of teams still fully manual, according to industry benchmark work and analyst commentary on balancing automation with human judgement.
The problem is what happens next. Benchmarks built specifically around creator workflows show that compiling creator data has become largely automatable while authenticity vetting remains the weak point. Models tested on full campaign pipelines let fraudulent creators onto final shortlists even when explicitly asked to screen for bought followers and engagement pods.
That is not an argument against the tools. It is an argument about where the human belongs. Automate the search. Automate the data pull. Keep judgment, brand safety, and the final call with a person. The most useful automation here produces evidence a human can check, rather than a score nobody can explain. Research on how recommendation systems surface options, on information overload in decision making, and on how ranking systems shape what people can find points the same way: more candidates do not help if the buyer cannot compare them.
What good discovery infrastructure looks like
Pull the threads together and a checklist falls out. It is not about which tool you buy; it is about which questions your process can answer on demand.
Public, structured creator profiles, so evaluation starts before anyone makes contact. Performance data that arrives with the profile rather than on request. Filters that work on audience composition first and creator popularity second. Authenticity signals attached to the profile, not buried in a separate audit. A record of past brand relationships, so shortlists reflect who has delivered rather than who has the loudest numbers. And one workflow that carries the same creator from discovery through negotiation, contract, and payment, because every handoff between disconnected tools is where information and accountability leak.
Infmap was built against that list: advanced search filters on the paid tier, public profiles for every user, a CRM per user, deal execution and deal management in one place, and payouts handled inside the platform. Research on platform and audience strategy, on how businesses build durable advantage, and on management practice suggests more teams will be shopping for exactly this once discovery costs become visible in a profit and loss statement.
How to structure the search so it compounds
The reason discovery keeps feeling like a fresh problem is that most processes leave nothing behind. A search happens, a deal closes or does not, and the learning evaporates. Fixing that is less about technology than about sequencing.
Start from the buyer problem, written in one paragraph. Turn it into the searches your buyer would actually run, then into the creators who already answer those searches. The study of search as a social system explains why this route surfaces creators who are genuinely useful rather than merely popular, and search marketing practice applies the same logic to audience discovery. Industry work on how people research purchases and on what converts that research into a decision is the input that keeps the list honest.
Then record where every name came from. Source tracking is what makes a creator list auditable and refreshable, and it is the difference between a pipeline and a folder of bookmarks.
Next, score on evidence before you contact anyone. Median recent performance, audience fit, format fit, sponsorship history, and one written sentence explaining why this creator fits this campaign. If you cannot write that sentence, cut the name. Writing the reason down is also the fastest way to discover that your criteria were never really defined, which is a diagnosis worth having early.
Finally, keep the record. What they were paid, what they delivered, how the sponsored content performed against their organic baseline, whether they hit the deadline, whether you would work with them again. That record is the asset. It is the only thing that makes the next search cheaper than this one, and it is what brand tracking research, trust measurement, and institutional confidence data all imply when they show how slowly durable reputation builds.
Measuring the search itself
One metric most teams never track is the cost of the search. How many hours went into building last month's shortlist? How many creators were contacted before one signed? What share of the candidates you evaluated were rejected, and for what reason?
Those are ordinary marketing operations numbers, and treating discovery as an operations problem changes how it gets managed. Analytics practice for conversion and funnel work, as documented by analytics practitioners and in campaign measurement documentation, translates directly: define the funnel stages, instrument each one, then find the step with the worst conversion and fix that first. Media intelligence tooling offers the same discipline for reputation and coverage, whether you look at social listening research, media analysis, content distribution data, media intelligence reporting, monitoring practice, or communications resources.
Set three numbers and review them quarterly. Time to shortlist. Cost per signed creator. Share of signed creators who came back for a second campaign. The third number is the one that tells you whether your discovery process is producing relationships or just transactions, and it is usually a humbling read the first time.
Benchmark context helps when someone asks whether the numbers are normal. Market sizing from aggregated marketing benchmarks, spend distribution from budget research and advertising outlook data, platform-level measurement from audience measurement and audio audience studies, and industry survey work from social platform research, trend analysis, and engagement benchmarks are all reasonable places to anchor a comparison, provided you read the methodology before quoting the number.
What this means for the next budget conversation
Discovery is usually invisible in budget discussions because it does not appear as a line item. It appears as a salary, as agency retainer hours, as campaign delays, and as the occasional creator partnership that produced nothing and nobody can explain why. Manual discovery, monitoring, and reporting are real operating costs that rarely get added up, and estimates of what those costs total for a mid-sized brand run into six figures over a few years when they are finally measured. Vendor total economic impact studies commissioned by platforms make the same arithmetic explicit, and independent analyst and trade reporting on marketing operations, from research blogs to consultancy insights, market analysis, and payments and commerce research, has been making the case that hidden process cost is the most common source of margin leakage.
The fix is not to search harder. It is to build the process so that each search leaves something behind: a vetted profile, a documented decision, a performance record attached to a name. That is what turns discovery from a recurring chore into a compounding asset, and it is what the marketplace research demonstrated when transparency shifted spending away from the biggest names toward creators who could simply be verified.
One more thing is worth keeping in view. Consumers now discover brands through roughly six different sources on average and no single channel reaches more than a third of them, according to global brand discovery research, the underlying market data, and audience research on digital behaviour. Creator partnerships sit inside that mix rather than above it. Which means the quality of your creator list is not a marketing detail. It is a direct input into how many of the right people ever hear about you at all.
The creators are there. The evidence is the part most teams have never built, and building it is the only part of this problem that compounds. Everything else is browsing.
If you want to see how a discovery-to-payout workflow looks in practice, you can create a free account on Infmap, and if you want to understand where the budget lands once the shortlist is built, read our breakdown of how to measure influencer marketing ROI and why most influencer marketing campaigns fail. For teams running many relationships at once, how agencies manage large creator rosters without chaos picks up where this article stops, and Infmap's pricing covers what the platform costs at each tier.