
“Don’t translate keywords. Research them natively.”
How many times have we given that advice in international SEO? I have. It’s still correct.
Before the research begins, someone must still choose the first phrase. That choice can determine which parts of a market become visible in the keyword data. A translated category can be accurate and locally natural while still missing the term people use to express eligibility, compliance, or professional practice.
In international SEO, the first seed is more than a starting point. It’s an assumption about how the market names itself.
The terms translation can miss
Start with your own market, in your own language, in your own tool: two lookups, 30 seconds.
Here’s where that first assumption starts to matter. In the same U.S. database, on the same day, “commercial drone operator training” returned nothing. “Drone pilot training” opened a 26,520-keyword set and surfaced “FAA Part 107.”
That was the difference between no keyword set and a usable route into the market.
The usual workflow looks sensible. A team defines a category in its home market, translates or localizes the name, enters that term into a local database, and expands what comes back.
The starting phrase is often treated as an administrative detail: the string sent to the translation vendor, the local freelancer, or the regional SEO team before the “real” research begins. Yet every expansion, cluster, and opportunity estimate that follows depends on that choice.
In regulated and qualification-led markets, the descriptive category may sit outside the vocabulary people use to express eligibility, compliance, or employability. The relevant query set may instead revolve around a license, professional card, exam, statutory title, or regulatory code.
Track, grow, and measure your visibility across Google, AI search, social, local, and every channel that influences buying decisions.
The institutional layer behind the category
I call these terms market artifacts: locally meaningful labels that determine eligibility, compliance, or professional practice but can’t necessarily be derived from the descriptive category name.
Part 107, EPA 608, NMLS, Series 7, TIP, RITE, and A1/A3 all belong to this layer, even though they represent different kinds of qualification or regulatory shorthand.
A translator can translate “commercial drone operator training.” Translation alone can’t tell them that the U.S. pathway is organized around Part 107 and the FAA Remote Pilot Certificate.
A researcher can describe a security-guard course in natural Spanish without yet knowing that Spain’s Law 5/2014 defines the TIP as the public credential used to accredit authorized private-security personnel.
The category can survive translation, while the qualification that controls entry never enters the set. Native research remains essential. But the familiar advice starts after the first seed has already been chosen.
- How do you discover the local term you need in order to begin discovering local terms?
That is the gap I wanted to test.
When a plausible description opens the wrong door
In the U.S. drone case, “commercial drone operator training” returned nothing in Semrush Broad Match or Related. Changing the seed to “drone pilot training” opened a set of 26,520 Related keywords. Within the first 1,000 rows used for the analysis, “FAA Part 107” appeared at rank 17 after deduplication.
The same contrast appeared in Spain. “curso de operador profesional de drones” returned no data. “curso de piloto de drones” returned 338 raw terms, 292 after normalization, with “AESA A1 A3” at rank 14.
These are operational alternatives, not laboratory-perfect minimal pairs. One is the sort of careful description an international team could construct before it understands the market. The other is a more conventional occupational label used within that market.
A team that starts from the wrong plain name may not receive a weaker keyword set. It may receive no set at all and interpret an entry failure as evidence about demand.
The same pattern appeared in English and Spanish. This wasn’t a Spanish-language peculiarity. It was a warning about the route used to enter an unfamiliar market.
What broke and what the test actually showed
The pattern first appeared in exploratory work across transport, healthcare, and workplace safety. Those examples generated the hypothesis, but they couldn’t confirm it.
I then designed a cleaner comparison in “food safety training.” It produced exactly the result I expected: The English expansion surfaced a ServSafe term while the Spanish expansion didn’t reach the local handler terminology.
A test that only produces findings isn’t a test
The comparison was invalid. Semrush defines Broad Match as returning keywords that contain the seed in different variations and word orders. The English result literally contained the English seed. The Spanish market term didn’t contain the Spanish one. I’d mistaken lexical eligibility for cross-market discovery.
So I discarded the result and rebuilt the test around a narrower question: Could a plausible descriptive seed recover a market artifact at all?
I selected eight new cases, four in Spain and four in the United States. For each one, I fixed a neutral descriptive seed, a local occupational seed, and a verified market artifact before retrieval. On Aug. 21, 2026, I ran 80 combinations across Semrush and DataForSEO, keeping lexical and discovery routes separate. The seeds, aliases, normalization rules, 1,000-row analysis limit, and decision thresholds were locked before the queries were run.
Similar recovery, different failure modes
The headline result was similar across the two discovery routes. The mechanism wasn’t.
In Semrush Related, the neutral seed recovered the artifact in two of seven observable cases. All five failures were empty result sets. Whenever the neutral seed produced a populated set, the artifact appeared in both cases and near the top. The dominant problem was entry failure: The seed never became a usable keyword neighborhood.
DataForSEO Keyword Ideas recovered the artifact in two of eight cases. This time, every neutral seed returned a populated set. Six neighborhoods existed but omitted the locked artifact within the first 1,000 canonical rows. That’s discovery failure: The research looked substantial, but the institutional layer was missing.

The providers also disagreed on which cases worked. Semrush recovered TIP and EPA 608 where DataForSEO didn’t. DataForSEO recovered Part 107 from the neutral seed where Semrush returned nothing, and recovered PER in a case Semrush couldn’t observe.
The purpose of the comparison was never to rank the tools. Every keyword platform has to transform a seed into a retrievable set through its own index and retrieval logic. The experiment shows why the provider and route belong in the methodology, rather than disappearing behind the final spreadsheet.
A keyword universe isn’t the market. It’s the market as constructed by a seed, an index, and a retrieval system.
The local occupational seed helped, but not reliably enough to become a rule. In Semrush, raw recovery rose from two cases to five. Under the prewritten criterion, however, only four of seven counted as improved.
TIP through higher neighborhood coverage, and A1/A3, Part 107, and NMLS through new recovery. EPA 608 moved from rank 27 to 14, but it was already recovered and its neighborhood coverage fell, so rank alone didn’t count. DataForSEO moved in the opposite direction, recovering one artifact from the local seed instead of two.
And I lost the sentence I most wanted to keep. “Use the market’s own term” is good practice, but the data doesn’t support treating it as a universal fix.
The test gave us a more useful conclusion. A keyword research workflow can fail at the entrance, fail inside a populated set, or vary according to the provider’s retrieval system. Those are different problems, and they require different responses.
From keyword list to market-entry map
The practical implication is larger than testing a few extra synonyms. International research shouldn’t begin with a translated keyword list. It should begin with a market-entry map: a record of the different vocabularies through which the same commercial space can be reached.
Before opening a keyword tool, I now write down four possible entry points:
- Neutral description: The phrase a competent outsider or translator would use to describe the category.
- Local occupational term: The name workers, customers, employers, or sales teams use in practice.
- Institutional artifact: The license, card, exam, title, or code used by the regulator or qualification body.
- Commercial or legacy term: A brand, historical credential, or market label that remains common even when the official terminology has changed.

The value lies not only in having four strings. It lies in recording where each one came from. A translator supplies one kind of evidence. Search logs, sales calls, and support tickets supply another. Job boards reveal how employers describe the role. Regulators and qualification bodies define what legally or professionally controls entry.
Read the regulator before the competitor. The EPA defines who needs Section 608 Technician Certification, while in Spain, AESA defines the A1/A3 training and examination route. Competitor pages can reveal local language, but they can also repeat the same descriptive assumptions as everyone else. Training providers and forums show which labels the market has commercialized or kept alive.
That provenance matters. Without it, a local-sounding term can be mistaken for the official one, an official term can be mistaken for the phrase customers use, and an obsolete label can enter a brief without anyone knowing why it’s there.
For each important seed, I now record the exact string, seed family, source, target market, provider and route, result-set size, whether a verified artifact was recovered, and the next action. That turns seed selection from an invisible assumption into an auditable part of the research.
The four families should be run separately at first. Combining them too early erases the evidence. You want to know which door opened the market, which one returned nothing, and which one generated a large but incomplete neighborhood. Only then does it make sense to consolidate the useful terms into a working keyword universe.
This also changes the role of the local reviewer. Their job is no longer limited to checking whether a translation sounds natural after the list has been built. They should help construct and validate the entry map before volume filters and clustering begin.
Diagnose the failure before measuring demand
A keyword tool can return the same visual signal for very different reasons. The research process needs a diagnostic layer before a zero or a thin set enters an opportunity model.
- Entry failure means the seed returns nothing. It doesn’t tell you that the market has no demand. The immediate response is to test another seed family, inspect the local SERP, check whether the occupation or qualification has another name, and run the same concept through a second provider. A zero shouldn’t become a market conclusion until those routes have been checked.
- Thin entry means the seed produces a handful of rows, often with weak or zero volume estimates. The category may genuinely be small, but it may also be badly represented by that string. This is where job listings, internal site search, paid-search query logs, customer language, and regulator terminology become especially valuable. The tool has given you too little evidence to stop, not enough evidence to deprioritize the market.
- Discovery failure is more deceptive. The tool returns hundreds or thousands of plausible keywords, so the work feels complete. Yet a verified license, exam, or qualification is absent. The next step is reverse seeding: start from the artifact, build its neighborhood, and measure what the descriptive route failed to expose.
- Provider disagreement means two systems construct different versions of the category. Averaging the outputs or merging them without labels hides the disagreement rather than resolving it. Keep the provider and retrieval route attached to every export, then inspect why each system reached or missed the terms that matter.
This diagnostic vocabulary is useful beyond regulated sectors. The “artifact” may be a product standard, a public funding scheme, a degree title, a procurement framework, a medical classification, a model number, or a legacy category name. The common feature is that the market organizes itself around a label an outsider wouldn’t naturally derive from the category description.
The decision rule I would adopt inside an SEO team is simple: no zero enters a market-sizing deck until at least one alternative seed family, one non-keyword source, and one second retrieval route have been checked. That’s a small process change with a potentially large effect on where budgets move.
Reverse seeding changes more than the keyword list
Reverse seeding is often described as another way to find keywords. Its more important value is that it can expose a different content system.
Take the drone example: “drone pilot training” can support a service or category page. “Part 107” opens questions about eligibility, the exam, preparation, certification, renewal, and operating rules. Those needs belong to the same commercial journey, but they’re not necessarily one page, and they’re not necessarily close in vocabulary.
The same applies to U.S. mortgage training: “mortgage loan originator training” describes the learning need. “NMLS” introduces licensing, approved education, registration, test preparation, renewal, and state-specific requirements. A study built only from the occupation may understate the size and complexity of the content opportunity even when it eventually finds the acronym.
This has three consequences for SEO architecture.
- Lexical distance doesn’t equal commercial distance. A regulator’s code and a plain category term may share almost no words while serving adjacent stages of the same decision. Clustering systems that rely too heavily on lexical similarity can separate commercially connected needs or merge terms that happen to look alike but belong to different journeys.
- The artifact often deserves its own hub. The occupational page can explain the service, course, or product. The artifact page can explain what the qualification is, who needs it, and how it relates to the offer. Supporting pages can address exams, renewals, eligibility, costs, documentation, and regional variations. Internal links then make the relationship explicit for users and search systems.
- The gap changes how opportunity is estimated. A neutral seed may return zero and contribute nothing to the forecast. The artifact-led neighborhood may reveal a substantial set of high-intent needs. If the opportunity model only includes the first route, the error isn’t confined to the keyword list. It affects the market score, the content budget, the launch sequence, and sometimes the decision to enter the country at all.
This is where the research becomes a business question. A seed isn’t simply a query we type into a tool. It can determine which part of the market becomes visible enough to receive investment.
A working method for international SEO teams
The study used 80 runs because I wanted to stress-test the idea. A normal project doesn’t need 80 runs. It does need a clearer sequence.
Before the tool, build the market-entry map. Start with the neutral description, then gather the occupational, institutional, and commercial vocabularies from people and sources close to the market. Ask a local practitioner to validate the distinctions, not merely the grammar.
Inside the tool, keep the routes separate. Run the seed families through lexical and discovery modes without merging the exports. Record whether each route produces an entry failure, thin entry, or populated neighborhood. If the providers disagree, keep that disagreement visible.
After the tool, reverse-seed the verified artifacts and compare the resulting needs with the descriptive set. The output shouldn’t be only a keyword list. It should include the missing journeys, proposed page types, internal-linking relationships, and any effect on market sizing.

The final research file should therefore carry more context than keyword, volume, and intent. At minimum, it should preserve seed family, seed provenance, target market, provider, retrieval route, diagnostic outcome, artifact recovery, and content implication.
That changes several common handoffs.
- Content briefs can state whether a page is built around the category, the occupation, or the qualification.
- Technical teams can see when region-specific URLs are justified by materially different institutional systems.
- Analysts can distinguish a genuine low-demand category from a weak entry point.
- Executives can see which market estimates are supported by multiple routes and which remain uncertain.
It also creates a better QA question: “Which local vocabularies did we test, where did they come from, and what part of the market did each one reveal?”
That’s harder to answer, but much closer to what international keyword research is supposed to do.
What this test doesn’t prove
This is a small practitioner study: eight held-out cases, two countries, two providers, and one execution date.
It measures tool output rather than buyer behavior. It can’t estimate market size, establish that a fixed percentage of regulated categories disappears from a neutral seed, prove that Spanish is harder than English, or turn the local occupational term into a universal solution.
It also measures the tool layer, not the search engine itself: Google’s own query rewriting and expansion may bridge part of this gap on the results page. But opportunity models are often built from tool exports before anyone inspects the SERP, and that’s where a zero can harden into a conclusion.
The seed construction is the largest limitation. The strings were locked before retrieval, but they weren’t independently proposed and rated by several international SEO practitioners.
Six of the 16 neutral and local seeds had no measurable direct volume. That’s consistent with the idea that a competent outsider can create a valid but market-unfamiliar description. It also means the exact figures should be read as a stress test, not as an estimate of how every experienced team would perform.
A stricter post-hoc observability check reduced the usable sample below the minimum set for a cross-case conclusion. The primary Semrush result survived the other sensitivity checks, but the sample remains small. DataForSEO also had an execution-order deviation: Its artifact controls were interleaved with some descriptive runs. The seeds and aliases were already locked and didn’t change, but that provider should be treated as robustness evidence rather than a procedurally perfect replication.
These limits narrow the claim. They don’t remove the practical lesson. Seed dependence is structural to keyword research, even though different providers expose it in different ways.
See where your brand appears, where it doesn’t, and exactly how to win more visibility across search, AI, local, social, and every channel that matters.
The first seed is a business assumption
International SEO has spent years improving what happens after the seed: local databases, native-language research, market-specific SERP analysis, regional architecture, and localized content.
All of that matters. That sits alongside a broader point I’ve made before: Spanish can’t be treated as one global market and cultural SEO begins with market-specific signals. This article moves the question one step upstream.
Before the first export, before the cluster, and before the content brief, someone chooses a string. That string expresses a theory about how the market names the category. The keyword tool then applies its own theory about what belongs near it. Neither theory appears in the finished spreadsheet.
Translation can preserve the category and still fail before research begins.
Misread the zero, and the mistake travels beyond the spreadsheet. A market can disappear from the opportunity model, budget can move to another country, and the content roadmap may never be built. The first assumption survives because the research looks finished.
Before you size a market, test whether you have entered it.
Run the descriptive name and the local occupational name side by side. Then run the license, card, exam, or qualification that actually controls the activity.
If one route returns nothing and another opens the category, you’ve learned something about the entry point before making a claim about demand.
Two lookups, 30 seconds, and then the interesting part.
