The Four Word Rule for Lower Amazon Ad Costs

The Four-Word Rule for Lower Amazon Ad Costs

TL;DR

  • Head terms run 1-2 words, mid-tail 2-3, long-tail 4 or more. The cost curve bends at four.
  • Score candidates as relevance × intent, both out of 5. “Blue yoga mat 1/4 inch” scores 20. “How to clean yoga mat” scores 2.
  • Harvest a search term into a manual campaign once it shows 2 or more orders and an ACoS below your target.
  • Start with 20 to 50 long-tail terms, not five hundred. The list is a working set, not an inventory.

Short version: the reason long-tail keywords cost less is not that they are cheap. It is that the people typing them have already decided.

A seller runs a campaign on “stapler.” Clicks arrive, spend arrives, orders do not, and the ACoS number climbs into territory that would be alarming if anyone were looking at it weekly. The campaign is not broken. It is bidding on a word that describes an object rather than a purchase.

“Heavy duty stapler for 100 pages” is the same product and a different transaction. The person typing it has specified a use case, a capacity and a durability requirement. They are not researching. They are checking whether you sell the thing they already pictured.

Where the Line Sits

The three bands are easy to state and worth stating precisely, because the boundary is where the economics change.

Head terms are one or two words. Mid-tail is two to three. Long-tail is four or more, and that is the band where the intent signal becomes strong enough to carry a campaign.

The reason is compression. Every additional word a searcher types removes a group of people who wanted something else. By the fourth word the remaining audience is small, specific and disproportionately made of buyers. Volume falls. Conversion rate rises faster than volume falls, which is the entire mechanism, and it is why a well-built long-tail set can run at a fraction of a head term’s ACoS on the same product.

Competition works the same way. A phrase with fewer than ten strong competitors is winnable in a timeframe that matters. A head term with the category’s incumbents on it is not, at any budget you are likely to have.

Six Families That Generate the Fourth Word

Long-tail terms are not found by staring at a list. They are generated by running a product through modifier families and seeing which combinations describe something real.

Six families do most of the work: size, material, pack quantity, feature, audience and compatibility. Applied to a mixing bowl they produce “3-pack stainless steel mixing bowls with lids.” Applied to cookware they produce “glass lid replacement for 10 inch skillet.” Both are four-plus words, both name a specific purchase, and neither is a phrase anyone would arrive at by sorting a keyword export by volume.

The full generation and validation workflow for long-tail keywords on Amazon runs this as six steps from build through harvest, which is worth following in order because the validation step is what stops the list becoming five hundred plausible phrases nobody has tested.

Score Before You Spend

Two axes, both out of five, multiplied. Relevance to what you actually sell, and purchase intent in the phrasing.

Term Relevance Intent Priority
Blue yoga mat 1/4 inch 5 4 20
How to clean yoga mat 2 1 2

 

The spread between 20 and 2 is the point. Both terms mention yoga mats. One is a purchase and one is a maintenance question from somebody who already owns yours. A volume-sorted list will frequently place the second above the first, because instructional queries are searched more often than transactional ones.

Multiplying rather than adding is deliberate. A term that scores 5 on relevance and 1 on intent lands at 5, not 6, and it deserves to. Low intent is not a small deduction, it is a disqualification.

Two terms about the same product, eighteen points apart. A volume-sorted list often ranks the wrong one first.
Two terms about the same product, eighteen points apart. A volume-sorted list often ranks the wrong one first.

The Harvest Rule

Discovery campaigns exist to produce evidence, and the evidence has a threshold: two or more orders with an ACoS below your target. A term that clears both gets promoted into a manual campaign where you control the bid. A term that does not stays in discovery or gets negated.

Two orders is not statistical significance and does not pretend to be. It is the point at which a term stops being a hypothesis. The alternative, waiting for a sample size that would satisfy a statistician, means spending several hundred dollars per term to find out, which is not a budget most sellers have across fifty terms.

Where the threshold genuinely matters is the ACoS half, because that number is meaningless without a margin to compare it against. This is the least glamorous part of the exercise and the one that decides whether any of it worked. The Small Business Administration’s guidance on managing business finances makes the point that cash and accrual accounting produce different pictures of the same business, which matters here specifically: marketplace fees are deducted at settlement, not at sale, so the margin you are measuring ACoS against is not the one in your pricing spreadsheet.

Three Layers, Three Jobs

The structure that holds all of this has three parts, and mixing them is the most common way a working keyword set stops working.

Discovery runs broad and automatic. Its output is search terms, not sales. Judge it on what it surfaces.

Control holds the terms that cleared the harvest rule, on exact match, at bids you set deliberately. This is where profit lives.

Scale takes the proven winners from Control and pushes budget at them. Only terms with a demonstrated conversion history belong here.

Keeping them separate means each layer can be judged on its own metric. Merged into one campaign, a discovery term’s poor ACoS drags the average of terms that are performing, and the reporting stops telling you anything actionable.

Start Small on Purpose

Twenty to fifty long-tail terms is the right opening set. It is small enough to fund a real test on each and large enough that a few will work.

The instinct to build a list of five hundred comes from treating keywords as an asset to accumulate. They are not. They are hypotheses to spend money testing, and a hypothesis you cannot afford to test is not on your list, it is on a spreadsheet.

Pick the twenty terms that score highest on relevance times intent, fund them properly, and run the harvest rule for a month. The set that survives is worth more than the five hundred you did not test.

Leave a Reply