Company Similarity Data for B2B ICP Expansion – August 2026

If your ICP expansion has started feeling like you’re just adding worse versions of your best customers, the seed accounts you’re matching against are probably the problem. Lookalike lists built on firmographics alone replicate the shape of your customer base without the substance. The accounts that close fast and expand well share behavioral patterns beyond company size alone. Use company similarity data to find those patterns and build a ranked account list your team can actually work.

TLDR:

  • Firmographic filters describe what a company is, not how it behaves: two firms can match on size and industry yet sit in completely different buying situations.
  • Build your seed list from your top 10 to 20 percent of closed-won deals, then capture tech stack, hiring activity, and news events at the time of purchase.
  • Layer buying signals (funding rounds, leadership hires, office expansions) on top of your lookalike list to rank accounts by readiness, beyond structural fit alone.
  • Score accounts on three inputs: similarity fit, signal stack depth, and strategic value. Teams using a scored ICP approach report win rates 20 to 40 percent higher.
  • PredictLeads Similar Companies returns up to 50 lookalikes per seed account across 18.8 million companies, with a plain-text similarity reason for the top 20 matches.

What ICP Expansion Actually Means in B2B Sales

ICP expansion is the deliberate process of identifying new account segments that share the buying behavior of your best existing customers, beyond their demographic profile alone. In practice, it means moving beyond the segment you have already saturated and finding adjacent markets where your product solves the same core problem for a structurally similar buyer. Done right, it extends your addressable pipeline without forcing your team to retool its pitch or its process.

What makes this hard in B2B sales is the gap between accounts that look like your ICP and accounts that act like it. Two companies can match on headcount, industry, and geography and still differ completely on buying motion, tech stack maturity, and the urgency behind a purchase decision. The first type fills a spreadsheet. The second type fills a pipeline. ICP expansion done well means finding the second type by matching on behavior as much as structure, which is why static firmographic filters consistently underperform similarity data built from technology detections, hiring activity, and news signals.

Why Firmographic Filters Cap Your Expansion

Filter your CRM by industry, headcount, revenue, and geography, and you get a list. What you do not get is any guarantee that those accounts behave like your best customers. That is the ceiling most GTM teams hit once they have worked through the obvious segment: the spreadsheet says these companies match, but pipeline says otherwise.

Firmographic overlap is a starting point, not a finish line. Two SaaS companies with 200 employees, $20 million in revenue, and headquarters in Austin can look like twins on paper and still diverge completely once you look at how they operate. One might be scaling a self-serve engine and hiring product-led growth marketers. The other might be stuck in a long enterprise sales motion, hiring solutions engineers, and running a tech stack built for six-month procurement cycles. Same filter, different buyer.

If your team has already exhausted the accounts that fit cleanly on size and industry, and the newer names you are adding convert at a lower rate, the problem usually is not effort. It is granularity. Firmographic filters describe what a company is. They say nothing about what it does, which tools it runs, or what roles it is actively hiring for, and those day-to-day details are what actually separate a buyer from a lookalike that just happens to share a NAICS code.

Here is what firmographic filters miss and behavioral signals catch:

  • Hiring direction: A company adding product-led growth marketers behaves nothing like one adding solutions engineers, even if both match on headcount and revenue.
  • Technology stack: The tools a company runs today tell you more about its buying readiness than the industry code it was assigned at incorporation.
  • Sales motion: A self-serve engine and a six-month procurement cycle produce two different buyers who happen to share a NAICS code.
  • Growth stage: Two companies at the same size can sit at completely different points in their growth cycle, one scaling fast, the other stalled.

This gap matters because the cost of targeting the wrong accounts is not small. Some sales-industry reports associate strong ICP fit with 2 to 3 times higher win rates and sales cycles that are 30 to 60 percent shorter, although results vary considerably by market and sales motion. Every account added to a list that looks right on a filter but fails on behavior is a rep working a longer cycle for worse odds. That is the case for moving past static filters toward similarity built on what companies actually do: the tools they run, the roles they hire for, and the events that signal a change in priorities. A company that matches on those dimensions actually behaves like your best customer; it does not merely resemble it on a spreadsheet.

What Company Similarity Data Is and How It Works

Company similarity data is a structured output that identifies which companies most closely resemble a given input company, based on behavioral and activity-level signals instead of shared demographic categories. Instead of asking whether two companies share a NAICS code or a headcount range, a similarity model draws on signals from across its datasets, including technologies, job openings, news events, connections, and company information. The output is a ranked list of lookalikes with a score and, in the case of PredictLeads Similar Companies, a plain-text reason explaining why each match surfaced. That reason field is what separates similarity data from a filtered list: a rep opening a new account sees “this company runs the same CRM stack and has been hiring sales operations roles for the past 60 days” and not merely a name with no context attached. The underlying signals update continuously, so a match that surfaces today reflects what a company is doing right now, not what it looked like when it was last categorized in a static database.

How to Define Your Seed Accounts Before You Expand

Before you run any similarity match, you need a clean answer to one question: which customers do you actually want more of? Skip this step and the algorithm will replicate your worst accounts alongside your best ones, because it has no way of knowing the difference unless you tell it.

Start with closed-won deals from the last 12 to 18 months. That window is recent enough to reflect your current product and positioning, but wide enough to give you a real sample. Pull the full list, then rank it across a few dimensions:

  • Revenue: How much the account is worth today, beyond what it was worth at signing.
  • Retention: Whether the account has renewed or is sitting in a risk bucket.
  • Expansion: Whether the account grew its contract value after the initial close.
  • Sales cycle length: Whether the deal moved fast or dragged through months of back and forth.

Take the top tier from that ranking, maybe the top 10 to 20 percent, and treat those as your seed accounts. These are the companies you want the similarity engine to replicate, not your entire customer base and not whatever logo happens to be easiest to pull from the CRM.

Once you have that shortlist, look past the deal metrics and capture what these accounts actually looked like from the outside at the point they became customers.

What to Capture About Each Seed Account

  • Technology stack: The tools an account had in place often explain why your product was a fit over a competitor’s.
  • Hiring activity: What a company was staffing up for around the time it bought matters. A company building out its sales team behaves differently from one investing in engineering, even if both became customers.
  • News events: A funding round or a leadership change that happened alongside the deal often explains the timing of the purchase, as much as the fit itself.
  • Product categories: What the company listed on its website tells you what kind of business it was building, which is a different question than the industry code it was assigned.

A seed list built this way gives the similarity engine something precise to work from. Feed in every logo in your CRM regardless of quality, and you get a lookalike list padded with noise, accounts that resemble your worst-fit customers just as much as your best ones.

Using Similar Companies to Expand Your ICP

Once your seed list is clean, run each account through Similar Companies and pull up to 50 lookalike companies per seed – 20 being the default. The matching logic draws on Technology Detections, Job Openings, News Events, and Connections simultaneously, so a match can surface because two companies share a tech stack, hire for the same roles, or appear on each other’s customer pages, and not simply because they share an industry code. That multi-dimensional matching is what separates the output from a filtered list: a company can look nothing like your best customer on headcount and revenue, yet still be an excellent fit because it runs the same tools and is staffing up the same way. For the top 20 matches per seed, each result includes a plain-text reason field explaining exactly why that company surfaced. Use that reason to skip generic outreach and open with context specific to the match. Across 10 to 20 seed accounts, you will typically surface several hundred candidates; the overlap between seeds is useful too, because an account that matches multiple top customers independently is a stronger signal than one that matches only one.

Layering Buying Signals on Your Lookalike List

A lookalike list tells you who to target. It does not tell you when to reach out, and timing is where most of the value in this exercise actually lives. A structural match sitting quietly with no recent activity is a name on a spreadsheet. A structural match that is hiring, expanding, or adopting new tech is an account with a reason to talk to you right now.

This is the layer most teams skip, and it shows up later as reps working a list cold. Spotio’s aggregated sales statistics report that 42 percent of sales reps lack sufficient information about a prospect before making first contact, according to Spotio’s sales statistics report, though this figure draws on aggregated third-party data rather than original research. A similarity score alone does not close that gap. It tells you a company resembles your best customer on structure, not that the company is doing anything today that makes it worth calling. Closing that gap means adding a second layer on top of the match: signals that show a company is actually moving.

A few signal types are worth checking for every account on your expansion list:

  • Hiring for relevant roles: A company adding headcount in a function tied to your product category, such as sales operations or a specific engineering discipline, is telling you something about where its priorities sit this quarter.
  • Technology changes: An account dropping a legacy tool or adding a new one in your category creates a pain point or a compatibility need that did not exist a month ago.
  • Office expansion: A company opening a new regional location is scaling operations, and scaling operations usually means new budget getting allocated somewhere.
  • Leadership hires: A new VP of sales or a new head of marketing often means a fresh set of vendor decisions within the first 90 to 180 days on the job.
  • Funding events: A recent funding round may indicate increased capacity to invest, although it does not guarantee an active buying process, and vendor evaluations are one early destination among many.

Structural similarity narrows the list. Buying signals tell you which names on that narrowed list are worth calling first. A company that matches your best customer on technology stack and hiring pattern, and has also opened a new regional office while adding sales operations headcount, is a different prospect than one that only clears the firmographic bar. The first account has a reason to be in-market. The second might be, eventually, but you have no evidence of it yet.

Stack enough of these signals and the list stops looking generic. It starts looking like a queue ranked by who is actually ready to have the conversation, the difference between a list a rep works in order and a list a rep works at random.

Scoring and Ranking ICP Expansion Accounts

A high similarity score tells you a company looks like your best customer. It says nothing about whether that company is ready to buy this quarter, next quarter, or at all. Treating similarity as a priority score is how a team ends up calling the most structurally perfect account on the list first, only to find no budget, no urgency, and no hiring activity anywhere near your product category. Fit and readiness are different questions, and your prioritization model needs to answer both.

A working model combines three inputs into one composite score, instead of leaning on any single number.

  • Similarity fit: The structural score itself, pulled from your lookalike match. This is the baseline, not the verdict.
  • Signal stack depth: How many buying signals are active on the account right now, and how recent they are. A company with three signals from the past 30 days outranks one with a single signal from six months ago.
  • Strategic value: Deal size potential, expansion room, and whether the logo carries weight in your market. A smaller account that could double at renewal deserves a different weight than one that caps out fast.

Weight each input based on what has actually driven wins for your team. If your best deals close on timing more than fit, weight the signal stack higher. If your sales cycle is long and logo selection matters more than speed, weight strategic value higher. The point is not to find a universal formula; it is to make the weighting explicit instead of leaving it to whichever rep feels like calling a name first.

Once every account has a composite score, sort into tiers instead of one long ranked list. A tiered view is easier to work and easier to review.

Tier

Composite profile

Action

Tier A

High similarity, active signals, strong strategic value

Fast, personalized outreach referencing the specific similarity reason

Tier B

Strong on one or two inputs, weak on the third

Slower nurture sequence, revisit as signals accumulate

Tier C

Low across most inputs

Stay on the list, recheck periodically for new signals

Tier C stays on the list because structural similarity usually changes more slowly than buying readiness, but both should be refreshed periodically, and an account with no signals today can surface one next month.

The payoff for doing this work shows up in the numbers. Factors.ai reports that teams using a documented, scored ICP see win rates 20 to 40 percent higher than teams using unscored lists, according to factors.ai’s ICP marketing guide. That gap is not about better accounts. It is the same expanded list, worked in an order that matches actual readiness instead of alphabetical order or gut feel.

Where Similarity Data Falls Short

Similarity matching is not a fit test. It tells you what a company looks like from the outside right now, not whether anyone inside that company has budget, buy-in, or the appetite to act this quarter. That gap does not disqualify the approach, but it does mean similarity output needs a second look before it becomes an account list. Three failure patterns show up often enough to name.

The first is circularity. If your seed accounts all come from one segment, say mid-market logistics companies that run the same three tools and hire the same three roles, the lookalike engine has nothing else to work from. It will hand you back a narrower version of the same segment, not a genuine expansion. This shows up most on teams with a young customer base, where the pool of closed-won deals is small and clustered by definition. The fix is to widen your seed set across whatever variation actually exists in your current customers, even if that variation feels small, and to treat a lookalike list that looks suspiciously homogeneous as a signal to check your inputs, not proof that the market is genuinely that narrow.

The second is competitor bleed. Similarity models frequently surface direct competitors of your customers as top matches, because the fastest way for a company to resemble your customer is to sell against them or sit in the same market. Sometimes that is useful: a competitor of your best customer, in the same niche, running similar tools, is a legitimate account if your product is not tied to their existing vendor relationship. If your CRM or competitive-intelligence data shows that an account is deeply committed to a competing solution, deprioritize it or route it into a longer-term displacement campaign rather than burning outreach cycles on it now.

The third is data thinness. A similarity score built from a narrow set of inputs, mostly headcount and industry code, will lean on firmographics because that is all it has, and it will produce matches that look right on paper but skip the layer that actually predicts fit: hiring direction, tech stack, recent news. A score with no visibility into what a company is doing day to day is a fancier firmographic filter dressed up as something more precise. Before trusting a similarity output, check what inputs actually feed the score. If it draws only from firmographic fields, treat it as a first pass and add job openings data, technology adoption signals, and news signals before you commit outreach time to the list.

How PredictLeads Supports ICP Expansion with Similar Companies Data

Similar Companies is where this approach becomes something you can run without stitching together five separate exports. PredictLeads covers 18.8 million companies in the dataset, returning up to 50 lookalikes per seed account, with a similarity reason included for the top 20 matches. That reason field is a plain-text explanation of why two companies matched, so a rep opening a new account starts with a line for outreach instead of a name with no context attached.

The matching logic does not lean on industry classification the way a basic lookalike tool does. It draws on signals from across the datasets PredictLeads maintains: technology detections, job openings, news events, connections, and company information. A match can surface because two companies share technologies, display similar hiring patterns, or appear in related customer, partner, or vendor relationships, which reflects a multi-dimensional signal instead of a single firmographic field dressed up as similarity.

Once you have a lookalike list, the next step is layering the buying signals covered earlier, and that is where pairing datasets matters. Similar Companies output pairs directly with:

  • Job Openings: covering 2.8 million or more companies with active hiring, so you can check whether a lookalike is staffing up for roles tied to your product category before you reach out.
  • News Events: 37 categorized event types across 2.5 million or more companies, covering triggers like funding rounds, leadership hires, and office expansions.

Running both against your lookalike list turns a static match into the tiered, signal-scored account list described earlier in this piece, without needing a separate tool for each layer.

All of it ships through API, flat files, webhooks, or a Model Context Protocol (MCP) server, so the delivery method matches how your team actually works. A RevOps practitioner enriching CRM records pulls through the API. A data engineer building a scoring pipeline for the whole account base takes flat files or webhooks instead. Either way, the underlying match and the reason behind it stay the same.


Ready to see this in your own data?

Get 100 free API requests/month – no credit card, no sales call.

Final Thoughts on Building a Stronger ICP Expansion Strategy

Getting ICP expansion right means starting with your best customers, not your whole CRM, and matching on behavior as much as structure. Add buying signals on top and you get a ranked queue your reps can actually work in order. That is the difference between a list and a pipeline. If you want to see what that looks like on your own seed accounts, start with 100 free API requests and run Similar Companies against your top closed-won deals.

Frequently Asked Questions

How do you build a seed account list for ICP expansion similar companies matching?

Start with closed-won deals from the last 12 to 18 months, rank them by revenue, retention, expansion, and sales cycle length, then take the top 10 to 20 percent as your seed set. Capture what each account looked like at the point of purchase: technology stack, active hiring roles, recent news events, and product categories listed on their website. Feeding a clean, ranked seed list into a similarity engine produces lookalikes that resemble your best customers, not your average ones.

Should I use firmographic filters or company similarity data for ICP expansion?

Firmographic filters describe what a company is on paper; similarity data built from technology detections, job openings, and news events reflects what a company is actually doing. Two SaaS companies at the same headcount and revenue can run completely different buying motions, and those differences only show up in behavioral signals, not in a NAICS code or headcount range.

What buying signals should I layer on top of a lookalike list before reaching out?

As a practical starting point, teams might prioritize accounts showing several recent and relevant signals rather than relying on one isolated event, which makes for a materially different prospect than one that clears only the similarity score.

How does PredictLeads Similar Companies data differ from basic lookalike tools?

PredictLeads Similar Companies covers 18.8 million companies and returns up to 50 lookalikes per seed account, with a plain-text similarity reason included for the top 20 matches. The matching logic draws on Technology Detections, Job Openings, News Events, and Connections simultaneously, so a match can surface because two companies share a tech stack, hire for the same roles, or appear on each other’s customer pages, and not simply because they share an industry code.

When does similarity scoring alone fail in B2B ICP expansion, and what should I add to the model?

Similarity scoring fails when your seed accounts come from a single narrow segment, because the engine has nothing else to work from and returns a contracted version of the same market instead of a genuine expansion. Pair the similarity score with signal stack depth, counting active buying signals from the past 30 days, and strategic value inputs like deal size and expansion potential, to build a composite score that ranks accounts by readiness instead of structural fit alone.

Is there a workflow for blending technographic data with hiring surges to spot ICP expansion opportunities?

Yes. Start with your Similar Companies list, then cross-reference each account against two signal layers: Technology Detections to see which tools a company has adopted or dropped recently, and Job Openings to see which roles it is actively staffing. A company that matches your best customer on tech stack and is also hiring for roles that signal spend in your category is a stronger prospect than one that clears only the similarity score. Use both signal types together in your composite score, weighting signal stack depth alongside structural fit, to rank accounts by readiness instead of fit alone.

Scroll to Top