Research
FSL begins as a language for seeing football differently.
Its first question is therefore not whether every FSL term can become a scientific construct.
It is simpler:
Does using the language help people see, describe or communicate football differently?
A useful distinction may remain qualitative. It may become part of football language without ever becoming a metric.
But some distinctions may invite further investigation.
- Can they be observed reliably?
- Can different people recognise them?
- Can they be operationalised?
- Can they be measured?
- Do they add anything beyond existing football concepts or data?
This page sets out ways of finding out.
Two levels of enquiry
FSL therefore has two related but different research questions.
1. Does the language work?
Does FSL help people:
- notice distinctions they might otherwise miss?
- describe football more precisely?
- communicate observations to other people?
- recognise recurring patterns?
- generate better questions?
This is the language and perception programme.
2. Can the distinctions survive empirical scrutiny?
If a distinction appears useful, a second question becomes possible.
- Can it be coded reliably?
- Can it be represented through observable variables?
- Can it be measured?
- Does it correspond to something beyond the act of describing it?
- Does it add information beyond established football concepts or measures?
This is the empirical programme.
A distinction needs to be worth seeing before it is worth measuring.
Four kinds of question
Different kinds of evidence answer different questions.
- Conceptual — Does the vocabulary make coherent distinctions? Do categories overlap unnecessarily? Does a proposed distinction add something to existing football language?
- Perceptual — Does using FSL change what people notice or how specifically they describe it? Can observers distinguish patterns more clearly? Can those distinctions be communicated?
- Methodological — Can different observers apply the definitions consistently? Can the categories be coded without relying on scoreline, reputation or hindsight?
- Empirical — Do the proposed patterns relate to observable features of football? Do they add information beyond existing concepts or measures?
These questions should not be treated as interchangeable. Reliability is not validity. Perceptual usefulness is not empirical proof. And empirical association does not automatically establish mechanism.
1. Does the language work?
The most direct test of FSL is also the simplest:
Does having the language change what people see?
A useful distinction should help an observer notice, describe or communicate something more precisely — not simply give them more football terminology.
Perceptual Programme
A controlled study could compare people trained in FSL with people using an alternative structured football vocabulary.
Participants could watch unfamiliar match passages and describe what they see. Independent, blinded judges could then assess:
- behavioural specificity
- temporal precision
- distinction between similar patterns
- collective rather than purely individual description
- ability to communicate what changed
- convergence between independent observers
The question is not who uses more football terminology.
It is who can communicate more useful distinctions about what is happening.
Recognition
Give observers unfamiliar passages and ask them to identify or describe the pattern they see. Test whether training changes recognition, specificity and consistency.
Transmission
One observer describes an unseen passage to another. The second observer then attempts to identify or reconstruct what the first observer saw.
This tests whether FSL vocabulary is communicable, rather than merely useful privately.
Convergence
Give independent observers the same passage without consultation. Examine whether their descriptions converge around similar distinctions.
Transmission and convergence are different. A language might help one person describe something precisely without necessarily producing agreement between people.
Transfer
A stronger test would use novel contexts. If the language is genuinely useful, its value should not depend entirely on having seen the original examples during training.
An active alternative vocabulary should be matched as closely as possible for training time, complexity, engagement and expectation.
2. Can the language be used reliably?
Once a distinction appears useful enough to investigate formally, the next question is whether independent observers can use it consistently.
The current vocabulary is therefore treated as a working coding system, not a finished taxonomy.
Coding system
Failure Patterns
Failure Patterns are provisionally coded as single-primary categories. The coder selects the pattern that best captures the dominant collective condition in the relevant passage.
Behavioural Patterns
Behavioural Patterns are currently treated as non-exclusive. More than one may be present in the same passage.
A team can, for example, Connect while also Driving, or Contain while beginning to Settle.
Transitions
Transitions are coded separately.
Where a transition is visible, the coder records:
- apparent starting pattern
- apparent destination
- approximate timing
- behavioural trigger or context
- confidence
- whether the change appears abrupt or gradual
The purpose is not to claim that one pattern caused another. It is to capture visible change in collective behaviour.
Reliability requirements
Reliability asks a basic question: Can independent observers use the same definition and arrive at reasonably similar judgements?
- use at least two coders where practical
- define coding rules before analysis
- keep coders independent
- avoid outcome, reputation or tactical shortcuts
- retain disagreements rather than silently resolving them
- use adjudication only after independent coding
- test adjacent categories where confusion is expected
- report uncertainty rather than forcing false precision
For Behavioural Patterns, agreement should account for the multi-select nature of the system.
Reliability is a prerequisite for interpreting later empirical tests. It is not evidence that a category is true.
Sample requirements
A useful research sample should contain enough variation to expose the strengths and weaknesses of the vocabulary.
For failure research in particular, include:
- tournament and high-stakes defeats
- late collapses
- goals conceded after sustained pressure
- visible periods of collective deterioration
- failure episodes within wins and draws
The last category matters. If failure is studied only when a team loses, it becomes difficult to separate collective failure from the outcome itself.
The sample should therefore contain failure without defeat, and defeat without assuming failure.
3. When the language meets data
Analytics gives a useful distinction a second life.
Once a way of seeing appears valuable, we can ask whether it can be translated into observable variables.
That translation may succeed.
It may reveal something new.
It may show that an existing metric already captures the distinction.
Or it may expose the distinction as too vague, unstable or dependent on interpretation.
All four outcomes are useful.
The aim is not to prove the language by turning every word into a number. The aim is to find out which distinctions, if any, survive that translation.
Data sources
- Match footage — for observing and coding collective behaviour.
- Event and tracking data — where available, for testing whether proposed distinctions relate to measurable features.
- Existing football metrics — to examine whether FSL adds anything beyond familiar measures.
- Expert language — to understand how experienced observers already describe collective failure and change.
- Controlled perceptual tasks — to test recognition, specificity, communication and convergence.
Different sources answer different questions. No single source is assumed to validate the language.
Empirical research phases
Phase 1 — Reliability & Feasibility
Can observers actually use the vocabulary? Can they identify the proposed patterns independently? Where do disagreements occur? Which definitions need revision?
Phase 2 — Descriptive & Analytical Tests
Once coding is sufficiently reliable, examine whether the patterns relate to observable match features.
Phase 3 — Player and Team-Level Questions
Some distinctions may eventually support questions about player roles, substitutions, development, team adaptation or longer-term performance.
These are later questions, not assumptions built into the language.
Analysis workflow
Observe → code → assess agreement → refine definitions → analyse relationships → compare with existing measures → test robustness
The order matters. A proposed relationship should not be treated as meaningful if the underlying category cannot first be identified with reasonable consistency.
Where quantitative tests are used, the analysis should distinguish between:
- association
- prediction
- incremental information
- discriminant validity
- possible mechanism
These are progressively stronger claims.
4. Candidate empirical questions
Once the language is sufficiently stable to test, several questions become possible. These are candidate hypotheses, not claims that FSL already makes.
- H1 — Failure and performance: Do periods coded as Collapse precede measurable deterioration in chance creation, territory or other indicators?
- H2 — Functional stability: Do behavioural patterns capture something not represented by possession or other conventional measures of control?
- H3 — Tactical distortion: Do different Failure Patterns produce recognisably different tactical changes?
- H4 — Transitions: Do observable changes in collective pattern occur before measurable changes in outcomes?
- H5 — Player contribution: Can player value partly consist in influencing collective patterns rather than only producing individual events?
- H6 — Opponent disruption: Can some behavioural patterns alter the opponent's available options or behaviour?
- H7 — Development: Does functional stability provide information about youth performance or development beyond conventional measures?
- H8 — Substitution: Do substitutions made during different Failure Patterns produce systematically different subsequent responses?
- H9 — Error cascades: Do destabilising events increase the probability of subsequent collective deterioration?
- H10 — Half-time pattern: Does behavioural condition at half-time contain information beyond the scoreline?
- H11 — Incremental validity: Do FSL distinctions add predictive or explanatory information beyond established football measures?
A positive result on any of these would be interesting. But a negative result would not automatically invalidate FSL as a language.
Expert language
Another natural research opportunity is to examine how expert football observers already explain collective failure.
A corpus of post-match analysis could examine what experienced observers talk about when explaining why a team lost control or conceded.
Descriptions could be classified by whether they concern:
- outcome
- event
- tactics
- individual action
- collective behaviour
- collective condition
- psychology or emotion
- intent
- opponent agency
- temporal sequence
This is not a test designed to make FSL win.
It could reveal that FSL names distinctions experts already perceive but describe inconsistently. It could reveal that FSL adds distinctions that existing language rarely captures. Or it could reveal that existing football language is already richer than FSL.
Any of those would be valuable findings.
What would make us change the language?
FSL should change when the language stops earning its keep.
That might happen if:
- observers interpret a category in incompatible ways
- a distinction does not improve description
- a category duplicates an existing concept
- users stop reaching for a term once its novelty disappears
- categories repeatedly collapse into one another
- definitions require arbitrary exceptions
- coding cannot achieve reasonable reliability
- operationalisation produces unstable or artificial measures
- empirical testing shows a distinction adds nothing
None of these would mean the project has failed.
They would mean: the language needs to change.
That is part of what the research is for.
What we know
FSL is currently a conceptual prototype.
The vocabulary is proposed rather than established. The underlying theoretical model is provisional. The coding system is subject to refinement. The empirical hypotheses are candidate questions. The perceptual programme remains to be tested.
No external review, audit or endorsement is implied unless explicitly documented.
The appropriate stance is therefore neither:
“These patterns are already proven.”
nor:
“They are only subjective, so nothing can be learned.”
It is:
Here is a language. Here are the distinctions it proposes. Now let's see what happens when people use it — and then what happens when we put those distinctions under pressure.
The open question
The deepest question is not whether FSL can eventually produce another metric.
It is whether better words can produce better distinctions.
If they can, some of those distinctions may become useful in coaching, analysis, discussion or research.
Some may survive measurement. Some may not. Some may turn out to be particularly useful ways of looking without becoming measurable constructs.
That is still worth finding out.