Clustering performance data reveals organizational and player roles
A position tells you where someone lines up. A win-loss record tells you what happened. Neither tells you what kind of team or player you're actually looking at — and that distinction turns out to matter more than either number.
Archetype, as we use it here, means a category defined by the shape of a team's or player's underlying performance data, rather than by a predefined label like "center" or "playoff team." An archetype is discovered, not assigned. That distinction is the entire premise of this post.
We reached this conclusion building a competitive-architecture model for the NHL. PlaymakerAI, a Swedish football analytics company, reached a strikingly similar conclusion working in soccer. Neither project referenced the other. Below is the reasoning that led us there, step by step, plus what PlaymakerAI's parallel work in a completely different sport tells us about why this approach works.
When we began classifying NHL franchises, the instinct was to ask "how good is this team?" That question only produces a ranking. Rankings compress everything down to one axis — better or worse — and in doing so, they throw away information about why a team is where it is and what kind of decision it should make next.
The better question, we found, was: what kind of organization is this, structurally? A team with a 58% Corsi rating that's never escaped the second round and a team with a 51% Corsi rating that's reached three conference finals are not the same organization wearing different records. They need different decisions made about them. A ranking can't see that. A classification can.
Definition — Corsi (CF%): the share of all 5-on-5 shot attempts a team generates out of the total attempts recorded in a game. It's a proxy for sustained puck possession and underlying system strength, independent of whether shots actually go in.
The methodological choice that matters most here is sequencing: explore before you classify. It's tempting to define categories first ("contender," "rebuilder") and then sort teams into them. That approach bakes in assumptions before you've tested them against the data, and it tends to miss categories that don't match anyone's existing mental model.
Instead, we started with five seasons of game-by-game data — Corsi, points, and playoff results across all 32 NHL franchises — and let the structure in that data surface on its own before assigning any names. Only after the groupings emerged did we label them.
Definition — clustering: an unsupervised method that groups observations (teams, players) by similarity across multiple variables simultaneously, without being told in advance what the groups should be. "Unsupervised" means the model isn't trained on pre-labeled examples; it finds structure in the data itself.
Two dimensions turned out to matter far more than any single stat: possession quality (CF%, a measure of system strength) and points trajectory — the direction of a team's results over the five-season window, not just its current level. A team holding steady at 90 points is a different organization than one that just fell from 100 to 90, even though a single-season snapshot would show them identically.
Layering in playoff depth — how far a team has advanced, and how consistently — as a third factor produced eleven distinct, data-grounded archetypes. A few examples:
Dynasty / System Team — elite possession, stable trajectory at a genuine peak, consistent deep playoff runs. The rarest classification: only two franchises qualified across all five seasons.
Goalie Rider / Clutch Performer — playoff success that outpaces what possession metrics would predict, often driven by a single elite performer. Sustainable only as long as that performer stays elite.
Regular Season Machine — strong possession and points, but a pattern of exits in the first or second round — evidence that something specific breaks down in series play that never shows up in the 82-game sample.
Stuck in Neutral — the most dangerous position on the map isn't the bottom; it's the middle. Not bad enough to justify a rebuild, not good enough to contend. The longer an organization sits there, the more expensive it becomes to leave in either direction.
The point isn't the specific labels. It's that the same current record can sit inside completely different archetypes, and each archetype implies a different next move.
This is where PlaymakerAI's work in soccer is a useful cross-check. Their team built what they call Playmaker Avatars — a machine learning model that uses football data to detect player roles, described in more descriptive terms than position alone, to make scouting easier. The reasoning behind it mirrors ours almost exactly, just applied to individual players instead of franchises: a center midfielder can play in very different styles, and the same is true of a full back or forward, so a single positional label hides more than it reveals.
Their avatars aren't tied to a formal position either. Each avatar has a common position it tends to match — for example, "The Worker" typically maps to a forward, and "The Physical" typically maps to a center back — but the category itself is defined by the shape of a player's statistical profile, the same way our archetypes are defined by the shape of a team's Corsi-and-trajectory profile rather than its raw record.
Two details from PlaymakerAI's approach are worth calling out because they apply directly to the NHL model too:
No player is a pure type. PlaymakerAI notes that every player is genuinely unique, with some matching a role closely and others reading as a mix of several avatars — which is exactly why a franchise's "archetype" is a classification, not a cage. A team can sit closest to Dynasty / System Team while still carrying traits of Regular Season Machine.
Visualization does the explaining a label can't. PlaymakerAI pairs its avatars with "avatar spiders" — radar-style charts that show how a player is actually being used and where their qualities might transfer to a different system. Our NHL work leans on the same instinct: the archetype map (CF% against points trajectory) is more useful shown than described, because the whole point is showing position within a structure, not just a category name.
Neither project was built to produce a leaderboard. Both were built to answer a harder, more useful question: given the shape of this team's or this player's underlying data, what should happen next?
In hockey, that means a Dynasty-tier front office and a Turnaround-tier front office facing the exact same question — "how do we get better?" — need opposite answers. One should protect the system that's already working; the other needs to build something that doesn't exist yet.
In soccer, that means a scout comparing two "central midfielders" on paper can instead compare their actual avatar profiles and know, before a single scouted match, whether the players are functionally interchangeable or fundamentally different in style.
That's the real case for archetypes over rankings or positions: they're discovered from the data's own structure, they hold up across sports that have nothing else in common, and they tell you not just where someone stands, but what to do about it.
© 2026 ORRO AI Genius, LLC. All rights reserved.