Resources

Glossary

Plain-English definitions for words that show up in the courses — grouped by career so you can skim the track you are on.

Data Science

Words you meet while cleaning tables, exploring A-level data, and writing a clear handover for a manager.

Dataset
A table of observations you analyse — rows are cases, columns are fields (for example subject entries by year and sex).
pandas
The main Python library for working with tables — load CSVs, filter rows, group, and summarise.
SQL
A language for asking questions of data stored in databases (SELECT, WHERE, GROUP BY).
Mean / median
Mean is the average; median is the middle value when sorted. Median is less pulled around by extreme outliers.
Chart
A visual summary (bar, line, box plot) that makes a pattern easier to brief than a raw table.
Handover note
A short written brief for the next person — what you found, how you got there, and what to watch out for.

Quantitative Finance

Market and portfolio language used when you measure risk, test a strategy, and write an investment research note.

Return
How much a price moved over a period, usually as a percentage (daily return, annualised return).
Volatility
How much returns jump around — higher volatility means a wider range of outcomes, not “more return”.
Correlation
How two assets tend to move together (from −1 opposite, to +1 lockstep). Diversification needs less-than-perfect correlation.
Moving average
A smoothed average of recent prices (e.g. 20-day) used to spot trend without reacting to every tick.
Backtest
Run a trading rule on historical prices to see how it would have performed — useful, but not a guarantee of the future.
Buy and hold
Own the asset for the whole period with no timing trades — the usual baseline a strategy must beat after costs.
Monte Carlo
Simulate many possible future paths using random draws from a model, then summarise the range of outcomes.
Efficient frontier
The set of portfolios that offer the best expected return for each level of risk (under a given model).

Actuarial Science

Pricing and risk words used when you model claims, set premiums, and brief a risk committee.

Frequency
How often claims happen (for example claims per policy-year).
Severity
How large a claim is when it occurs (claim amount distribution).
Pure premium
Expected loss cost ≈ frequency × severity — before expenses, profit, or loading.
Loading
Extra amount added to pure premium for expenses, profit, or uncertainty.
Relativity
How much riskier (or safer) one segment is versus a base — e.g. an age-band factor of 1.3 means 30% higher expected cost.
IBNR
Incurred but not reported — claims that have happened but are not yet on the books; estimated with techniques such as chain-ladder.
Reserve
Money set aside for unpaid claims (reported and IBNR).
Reinsurance
Insurance for insurers — transfer part of large losses to another carrier for a premium.
Stress test
Replay the book under a worse scenario (flood, longevity) to see whether capital and decisions still hold.

Machine Learning

Core modelling language for features, training, and honest evaluation before you specialise into computer vision.

Feature
An input measurement the model uses (hours studied, pixel values, a column in a table).
Label
The answer you want the model to predict (passed/failed, habitat class, clothing category).
Training set
Examples the model learns from — weights are updated on these.
Validation set
Held-aside examples used while training to check progress and tune choices (epochs, hyperparameters) — not the final exam.
Test set
Examples never used for training or tuning — evaluate once at the end for an honest score.
Epoch
One full pass through the training set.
Accuracy
Percentage of examples labelled correctly on a given set.
Overfitting
The model memorises training examples but fails on new ones — train score high, validation flat or falling.
Classifier
A model that assigns each example to a category (spam/not spam, habitat class).
Neural network
Layers of simple units (neurons) that learn weights from data; deeper stacks can represent more complex patterns.

Computer Vision (Analyst)

Map-reading language for analysts who interpret model outputs and write briefs rangers can act on — without building the model.

Land-cover class
A habitat or surface label on the map (forest, grassland, water, urban, and so on).
Tile / patch
A small square cut from a larger image that the model scores as one class.
Confidence
How sure the model is about a prediction — high confidence is not the same as being right; always check the map.
Confusion matrix
Table of true labels (rows) versus predicted labels (columns). Off-diagonal cells are the mistakes that matter in the field.
Limitation
An honest bound on what the map can support (season, cloud, rare classes) — required in briefs even when accuracy looks high.
Park brief
A short note for rangers or ecologists — what the map shows, where it is weak, and what action (if any) follows.

Computer Vision (Developer)

Developer vocabulary for building and training convolutional models on satellite tiles and reporting results.

Convolution / filter
Slide a small grid of numbers over an image, multiply and sum — highlights edges or textures.
Pooling
Shrink a feature map by keeping the strongest signal in each small block.
Batch size
How many images the model looks at before updating its weights once.
Augmentation
Random flips, rotations or brightness changes during training so the model sees variation without new photos.
Backbone
A pre-trained network (e.g. MobileNetV2) that already knows edges and textures; you add a small head for your land classes.
Transfer learning
Reuse a model trained on millions of general photos, then retrain only the top for your habitats.
Softmax
Final layer that turns raw scores into class probabilities that add up to 100%.
Stratify
When splitting data, keep the same mix of classes in each part (e.g. same proportion of Forest).
Confusion matrix
Table of correct labels (rows) vs predicted labels (columns). Off-diagonal cells are mistakes.