Education · textbook
Six parts, 24 chapters, 79 methods — neural, symbolic, and the hybrids that hold both. Each method is one paragraph you can read in fifteen seconds, a note on what to read first, and a link to the paper it came from. Each chapter points at several books and courses covering the same ground, so you can use whichever you already have.
learns the function from examples
Parameters fitted by gradient descent over data. Strong where the rule is easy to demonstrate and hard to write down: perception, language, anything with a fuzzy boundary.
The cost: the learned rule is distributed across billions of weights, so "why did it answer that" has no short answer, and out-of-distribution behaviour is discovered rather than derived.
executes the rule you wrote
Explicit structures — logic, constraints, graphs — manipulated by rules that preserve meaning. Strong where the rule is known, must be audited, or must never be violated.
The cost: somebody has to write the rules, they are brittle at the edges of their vocabulary, and the world rarely arrives pre-formalised.
learns the parts, checks the whole
Perception learned, structure enforced. A network proposes; a solver, grammar or knowledge base disposes — or the symbolic side supplies priors the network would need far more data to infer.
The cost: two systems to build, an interface between them that is usually the hard part, and far fewer people who have shipped one.
Six parts, 24 chapters, 79 methods. The order is a reading order: each part assumes the one before it, and every method lists what to read first if you arrive in the middle.
The oldest working idea in the field: if you can say what a solution looks like and what a move is, you can look for one. Everything here is a different answer to the question of where to look next.
Exhaustive search, and the cost of having no idea which way is better.
readArtificial Intelligence: Foundations of Computational AgentsPoole and MackworthArtificial Intelligence: A Modern ApproachRussell and Norvig
watchCS 188: Introduction to Artificial IntelligenceUC BerkeleyCS50's Introduction to Artificial Intelligence with PythonHarvard
Explore a state space by expanding nodes in a fixed order, with no estimate of how far the goal is. The baseline every other search is measured against, and the vocabulary — states, actions, frontier, path cost — that the rest of the field inherited.
Order the frontier by cost so far plus an estimate of cost remaining; with an admissible estimate the first solution found is optimal. The result that made search practical, and the reason so much of AI is about finding a good heuristic rather than a better search.
Spaces too large to enumerate, and worlds that change while you search them.
readArtificial Intelligence: Foundations of Computational AgentsPoole and Mackworth
watchCS 188: Introduction to Artificial IntelligenceUC Berkeley
Keep one state, or a few, and move to a better neighbour; accept the occasional worse move to escape a local optimum. Gives up completeness for memory, which is the right trade when the state space cannot be enumerated and any good solution will do.
Plan when actions have several possible outcomes or the state is only partly visible, producing a contingent plan rather than a sequence. Where planning stops being a path and becomes a policy — the step that separates a route from an agent.
Someone else is choosing half the moves, and they are not helping.
readArtificial Intelligence: A Modern ApproachRussell and Norvig
watchCS 188: Introduction to Artificial IntelligenceUC BerkeleyGame TheoryCoursera
Search a tree in which another agent chooses alternate moves against you, pruning branches that cannot affect the outcome. Alpha-beta prunes without changing the answer, which is the same shape of argument as speculative decoding: do less work, return an identical result.
Build the tree asymmetrically by sampling playouts, spending expansion where the estimates are most uncertain. Needs no evaluation function to start, which is why it took games where nobody could write one — and why it keeps reappearing wherever a model can score its own rollouts.
Stop describing the path and start describing the answer; let a solver find it.
readHandbook of Practical Logic and Automated ReasoningHarrisonArtificial Intelligence: Foundations of Computational AgentsPoole and Mackworth
watchCS 188: Introduction to Artificial IntelligenceUC Berkeley
Variables, domains, constraints; solved by propagation and backtracking. The natural formulation for scheduling, allocation and configuration — including, in this domain, fitting a model set onto hardware.
Decide whether a propositional formula can be satisfied. NP-complete and, thanks to conflict-driven clause learning, routinely solved at industrial scale — one of computing's most useful gaps between worst case and practice.
SAT plus theories — arithmetic, arrays, bitvectors — so constraints can be stated in the vocabulary of the problem. The engine under most program verification.
Search where the actions are the moves and the world is the board.
readActing, Planning, and LearningGhallab, Nau and TraversoModern Robotics: Mechanics, Planning, and ControlLynch and Park
watchModern Robotics: Mechanics, Planning, and ControlCoursera
Search for a sequence of actions transforming an initial state into a goal, given operators with preconditions and effects. Plans can be inspected and justified before anything acts, which is why autonomy in regulated settings keeps returning to it.
Plan a collision-free path through configuration space, while estimating where the robot actually is. Where perception, state estimation and planning stop being separate chapters and have to run in the same loop, on the same hardware, in real time.
The only part of the field where “prove it” is a literal instruction. Structures whose meaning is fixed by definition rather than inferred from data — which is why they compose, why they can be checked, and why they break the moment the world exceeds their vocabulary.
Two languages, and what each one can and cannot express.
readHandbook of Practical Logic and Automated ReasoningHarrisonClassical LogicStanford Encyclopedia of Philosophy
watchIntroduction to LogicCourseraCS50's Introduction to Artificial Intelligence with PythonHarvard
Atoms and connectives, no variables. Decidable, and the target language for an enormous amount of practical reasoning once a problem is encoded into it.
Adds quantifiers, variables and relations, so you can say something about all of a class. Expressive enough for most modelling, and only semi-decidable — a proof search may simply not return.
How a conclusion is actually derived, step by mechanical step.
readHandbook of Practical Logic and Automated ReasoningHarrisonArtificial Intelligence: Foundations of Computational AgentsPoole and Mackworth
watchIntroduction to LogicCoursera
Find the substitution making two expressions identical. The primitive operation beneath logic programming and type inference alike.
A single inference rule, complete for refutation: assume the negation and derive a contradiction. Made automated theorem proving a mechanical procedure rather than an art.
Robinson, A Machine-Oriented Logic Based on the Resolution Principle
Run the rules from what is known toward what follows, or from a goal back toward what would establish it. Data-driven monitoring versus goal-driven diagnosis.
Committing to what exists, so that a machine and an auditor can agree on it.
readOWL 2 Web Ontology Language PrimerW3CA Description Logic PrimerKroetzsch, Simancik and HorrocksModal LogicStanford Encyclopedia of Philosophy
watchCS 520: Knowledge GraphsStanford
Deliberately restricted fragments chosen so inference stays decidable and usually tractable. The formal basis of OWL and of most production ontologies.
Modal, temporal, deontic and defeasible systems for necessity, time, obligation and conclusions that later evidence can withdraw. Temporal logic is what model checkers verify against.
An explicit, shared specification of what exists in a domain and how it relates. Makes disagreement visible: two systems that cannot agree on an ontology did not agree before, they just had not noticed.
Facts as typed edges over identified entities, with a schema expressive enough to infer edges nobody wrote down. Practical strength is provenance: each triple can carry where it came from.
Condition-action rules over working memory, matched efficiently by algorithms such as Rete. Still the backbone of policy engines, where the rules must be read by the people accountable for them.
Grammar before statistics: the tradition neural NLP replaced, and did not entirely.
readSpeech and Language ProcessingJurafsky and Martin
watchCS 224N: Natural Language Processing with Deep LearningStanfordNatural Language Processing SpecializationCoursera
Recover structure from a string according to an explicit grammar. Displaced for understanding language and indispensable for constraining it: constrained decoding is a parser attached to a sampler.
A ladder of grammar classes, each able to describe more structure than the last and each costing more to recognise. The reason a regular expression cannot balance brackets and a parser can. It also fixes what a constrained decoder can enforce cheaply: regular constraints are free at sampling time, context-free ones are not.
Map a sentence to a formal expression that can be executed rather than merely interpreted. The oldest working answer to grounding: if the output is a query or a program, it either runs and returns something checkable or it does not. Text-to-SQL is this with a modern face.
Zettlemoyer and Collins, Learning to Map Sentences to Logical Form
Nothing above survives contact with a world you only partly observe. This part is what replaces certainty: degrees of belief that compose correctly, and a rule for acting on them.
Why probability, and not something more convenient.
readProbabilistic Machine LearningMurphyMathematics for Machine LearningDeisenroth, Faisal and Ong
watchProbabilistic Graphical Models 1: RepresentationCoursera
Degrees of belief that obey the probability axioms, updated by conditioning on evidence. The reason a system can say how sure it is in a way that composes — and the discipline a confidence score invented for a string match does not have.
Update a belief by weighting a prior with how well each hypothesis predicted what you saw. One line of algebra, and the only part of it that scales is the independence you are willing to assert. Every graphical model below is a bookkeeping device for those assertions.
Whether a stated confidence of 0.8 is right about eight times in ten. Accuracy and calibration are different properties, and modern networks improved one while getting worse at the other. It matters wherever a number is handed to a human or a threshold: an uncalibrated 0.95 is not a probability, it is a score.
Guo et al., On Calibration of Modern Neural Networks
Independence stated as structure, which makes it an auditable claim.
readProbabilistic Machine LearningMurphyArtificial Intelligence: Foundations of Computational AgentsPoole and Mackworth
watchProbabilistic Graphical Models 1: RepresentationCourseraProbabilistic Graphical Models 2: InferenceCoursera
A directed graph of conditional dependencies, with inference by message passing. Structure states what is assumed independent — an auditable modelling claim, not a fitted side effect.
Sum out the variables you did not ask about, one at a time, reusing the intermediate factors. Exact and often cheap, until the graph has a wide junction tree — at which point the cost is not a matter of implementation and you switch to sampling. Knowing which case you are in is the useful part.
Each node tells its neighbours what it believes; on a tree the messages converge to the exact answer. Run on a graph with loops it has no such guarantee and often works anyway, which is either an embarrassment or a tool depending on whether you have to certify the result.
Murphy, Weiss and Jordan, Loopy Belief Propagation for Approximate Inference
Answer a question about a distribution by drawing from it instead of summing over it. Trades an exact answer you cannot compute for an approximate one you can, with an error that shrinks in a way you can reason about. Whether the chain has mixed is not something the sampler will tell you.
Neal, Probabilistic Inference Using Markov Chain Monte Carlo Methods
Belief that has to survive the next observation, and the one after.
readProbabilistic Machine LearningMurphyArtificial Intelligence: A Modern ApproachRussell and Norvig
watchProbabilistic Graphical Models 2: InferenceCoursera
Track a hidden state through noisy observations by alternately predicting forward and correcting on evidence. Filtering, smoothing and prediction are the same machinery pointed at now, the past and the future; still what flies aircraft and locates robots.
A discrete hidden state that moves by a transition table and emits an observation at each step. Three questions, three algorithms: the likelihood of a sequence, the most likely path, and the parameters. It ran speech recognition for two decades and still runs alignment problems where the state really is discrete.
Rabiner, A Tutorial on Hidden Markov Models and Selected Applications in Speech Recognition
The same predict-and-correct loop when the state is continuous, the dynamics linear and the noise Gaussian — which makes both steps closed-form. Optimal under assumptions almost nothing satisfies exactly, and close enough often enough that it navigates aircraft. The extended and unscented variants buy non-linearity back at the cost of the optimality proof.
Represent the belief as a cloud of weighted samples, move them by the dynamics, and reweight on the evidence. Drops every assumption the Kalman filter needs and pays in samples. The failure mode is particle depletion: the cloud collapses onto one hypothesis and the filter becomes confidently wrong.
Rules that hold usually — the seam where the two traditions first met.
readNeurosymbolic AI: The 3rd Waved'Avila Garcez and Lamb
watchProbabilistic Graphical Models 1: RepresentationCoursera
Weighted first-order rules, so a formula can be usually-true instead of always-true. The most direct bridge from logic toward learning, and the ancestry of much neurosymbolic work.
Write the generative story as a program; let the language do inference over the values that could have produced the data. The model becomes source code — reviewable, version-controlled, and separable from the inference algorithm that runs it. That separation is the whole claim, and it is also where the performance goes.
van de Meent et al., An Introduction to Probabilistic Programming
From what is true to what to do, including when others are deciding too.
readReinforcement Learning: An IntroductionSutton and BartoArtificial Intelligence: A Modern ApproachRussell and Norvig
watchCS 188: Introduction to Artificial IntelligenceUC BerkeleyReinforcement LearningCourseraGame TheoryCoursera
Represent preferences as a utility function and choose the action with the highest expected utility. Separates what a system wants from what it believes, so the two can be argued about independently — which is the whole basis of stating an objective a reviewer can challenge.
States, actions, transition probabilities and rewards; solved for a policy that maximises expected return. The formalism underneath reinforcement learning, and the reason an RL result can be stated as a claim about a policy rather than about a training run.
An MDP where the state is not observed directly, so the agent acts on a belief distribution instead. The honest model of almost every real deployment, and computationally brutal — which is why most fielded systems approximate it and should say which approximation.
Decisions where the outcome depends on other agents choosing at the same time, analysed through equilibria and the incentives a mechanism creates. Becomes practical rather than academic the moment several models negotiate, bid or delegate to each other.
Fitting the rule instead of writing it. Nearly all of it is one idea applied repeatedly — define a differentiable function, define a loss, follow the gradient — but the part that came before gradients still matters, because it produced models you can read.
Models whose output is a rule, a tree or a program rather than a weight matrix.
readMachine LearningMitchellAn Introduction to Statistical LearningJames, Witten, Hastie and TibshiraniThe Elements of Statistical LearningHastie, Tibshirani and Friedman
watchSupervised Machine Learning: Regression and ClassificationCoursera
Recursive partitioning into readable tests. Ensembles of them remain hard to beat on tabular data, at the cost of the readability that motivated a single tree.
Learn logic programs from examples plus background knowledge. Learns from a handful of instances and returns a rule a domain expert can read and correct — the opposite trade to a network.
Muggleton, Inductive Logic Programming
Fit a model with hidden variables by alternating between inferring them and re-estimating parameters. Learns structure nobody labelled, and converges to a local optimum — so the initialisation is part of the result, not a detail.
Search for a program satisfying a specification or examples. Generalises from very little data because the hypothesis space carries the structure of a programming language.
The machinery every modern model is built out of.
readDeep LearningGoodfellow, Bengio and CourvilleDive into Deep LearningZhang, Lipton, Li and SmolaUnderstanding Deep LearningPrinceMathematics for Machine LearningDeisenroth, Faisal and Ong
watchDeep Learning SpecializationCourseraMIT 6.S191: Introduction to Deep LearningMITNeural Networks: Zero to HeroAndrej KarpathyImproving Deep Neural NetworksCoursera
Stacked affine maps with a nonlinearity between them. Universal in principle, inefficient in practice — every later architecture is a statement about which connections are worth not having.
The chain rule applied in reverse across a computation graph. Not a learning theory — an efficient way to get the derivative, which is why it outlived every theory it was attached to.
Rumelhart, Hinton and Williams, Learning representations by back-propagating errors
Estimate the gradient on a minibatch and step; adaptive methods rescale per parameter. A theoretical guarantee traded for the ability to train something enormous before the funding runs out.
Constraints that cost training accuracy to buy generalisation. Each encodes a belief about what the data would have looked like had there been more of it.
Keep activations and gradients in a workable range, and give the signal a path that skips the depth. Between them, the reason networks with hundreds of layers train at all.
Getting the world into a vector space without being told what the labels are.
readDeep LearningGoodfellow, Bengio and CourvilleUnderstanding Deep LearningPrince
watchDeep Learning SpecializationCourseraCS 224N: Natural Language Processing with Deep LearningStanford
Meaning as position in a vector space, so that similarity becomes geometry. The move that lets discrete things — words, users, molecules — enter a differentiable system at all.
Compress then reconstruct; whatever survives the bottleneck is what the data was about. The ancestor of most self-supervised objectives.
Learn by pulling matching pairs together and pushing mismatched ones apart. Turns unlabelled data into supervision, which is why it underlies modern vision-language models.
What changes between eras is the shape of the function and how cheaply it can be made wide.
readDive into Deep LearningZhang, Lipton, Li and SmolaUnderstanding Deep LearningPrinceComputer Vision: Algorithms and ApplicationsSzeliski
watchCS 231n: Deep Learning for Computer VisionStanfordConvolutional Neural NetworksCourseraCS 224N: Natural Language Processing with Deep LearningStanfordSequence ModelsCoursera
Weight sharing across space, encoding the prior that a feature means the same thing wherever it appears. A symbolic assumption, compiled into the architecture.
State carried across a sequence, with gates to keep gradients from vanishing. Displaced by attention for long context, still competitive where memory is tight and streams are endless.
Hochreiter and Schmidhuber, Long Short-Term Memory
Every position attends to every other, so the model chooses its own dependencies rather than inheriting them from the architecture. Quadratic in sequence length, which is the tax the KV cache pays at serving time.
Vaswani et al., Attention Is All You Need
Sequence mixing through a linear recurrence with structured state, giving near-linear scaling in length. The live argument against attention's quadratic cost.
Many parameter blocks, few active per token. Total parameters set what must be resident in memory; active parameters set how fast it decodes — which is why this registry records both, and why treating an MoE as dense understates its speed by the sparsity factor.
Learn to reverse a noising process, generating by repeated denoising. Dominant for images, audio and video; increasingly argued for over text.
No labels, only consequences — and, lately, preferences.
readReinforcement Learning: An IntroductionSutton and BartoSpinning Up in Deep RLOpenAI
watchReinforcement LearningCourseraSample-based Learning MethodsCourseraGenerative AI with Large Language ModelsCoursera
Learn a policy from reward rather than from labelled answers, by acting and observing what follows. Solves an MDP without being handed its transition probabilities. Its costs are the honest ones to state: sample hunger, reward specification, and behaviour that is only as safe as the environment it was allowed to explore.
Fit the model to comparisons rather than to a single target, when quality is easier to rank than to specify. Direct methods optimise the preference objective without training a separate reward model.
The part no textbook covered until recently and every deployment runs into immediately. A model you cannot fit, adapt or serve at a price you can pay is a result, not a system.
Making weights that are not yours behave as though they were, without retraining from scratch.
readDive into Deep LearningZhang, Lipton, Li and Smola
watchGenerative AI with Large Language ModelsCourseraAI EngineerYouTube
Adjust a pretrained model on a narrower distribution; low-rank adapters train a small factored update instead of the full weights. What makes on-premises specialisation affordable.
Store and compute weights at reduced precision. The single most consequential deployment decision: it sets whether a model fits the hardware at all, and every figure in this registry's fit arithmetic is a function of it.
Train a small model on a large one's output distribution. Often the only route to an edge-deployable model whose behaviour still resembles the one that was evaluated.
Hinton, Vinyals and Dean, Distilling the Knowledge in a Neural Network
Where the latency and the bill actually come from.
watchAI EngineerYouTube
Attention keys and values retained across decoding steps so each token is not recomputed. Independent of the weights and linear in context and concurrency, which is why it dominates memory long before the weights do.
Manage the KV cache in fixed pages, as an operating system manages memory, and admit new requests mid-flight. The difference between a demo and a service.
Kwon et al., Efficient Memory Management for LLM Serving with PagedAttention
A small model drafts, the large one verifies in parallel and accepts the prefix that matches. Faster output with identical distribution — a rare free lunch, and a genuinely symbolic move: propose, then check.
Leviathan, Kalman and Matias, Fast Inference from Transformers via Speculative Decoding
The two halves fail in opposite directions: networks are ungrounded and unaccountable, symbolic systems are brittle and hand-fed. Combining them is obvious and has been attempted for decades; what changed is that the neural half is now strong enough to be worth constraining.
Four seams, and what crosses each one.
readNeurosymbolic AI: The 3rd Waved'Avila Garcez and Lamb
watchAI EngineerYouTube
The network perceives; its output becomes symbols a solver reasons over. Perception is learned, reasoning is auditable. The interface is the hard part: a symbol asserted with false confidence is worse than a low score, because downstream logic treats it as fact.
Known structure supplied as architecture, loss or constraint, so the model is not made to rediscover it from data. Convolution is the oldest example; graph networks, physics-informed losses and typed decoding are the modern ones.
Relax logical operators into continuous functions so rules can be trained through. Tight integration and one gradient; the relaxation is also where guarantees leak, since an almost-satisfied constraint is not a satisfied one.
Restrict generation to a grammar or schema so output is valid by construction rather than by inspection. The cheapest neurosymbolic technique in production, and the reason structured output can be relied upon.
The model writes a query, program or plan; a deterministic engine executes it; the result returns. Arithmetic, retrieval and search move to systems that are correct by construction, and the executed artefact is a record of what was actually done.
Named systems that made a particular seam work, and are specific enough to argue with.
readNeurosymbolic AI: The 3rd Waved'Avila Garcez and Lamb
Probabilistic logic programming where predicate probabilities come from neural networks, trained end to end through the logic. The clearest demonstration that the two halves can share one objective.
First-order formulae interpreted over real-valued tensors, so logical satisfaction becomes a differentiable objective. Maximised alongside a learning loss.
Badreddine et al., Logic Tensor Networks
Learn visual concepts and a symbolic executor jointly from question-answer supervision, without ground-truth programs. Sample efficiency and compositional generalisation are the headline results.
Use background knowledge to guess the latent symbols that would explain an observation, then train perception against those guesses. Learns when labels are absent but the rules are known.
The chapter this company exists for: a check outside the model that it cannot talk its way past.
readNeurosymbolic AI: The 3rd Waved'Avila Garcez and LambPatterns, Predictions, and ActionsHardt and Recht
watchArtificial Intelligence: Ethics and Societal ChallengesCoursera
A symbolic monitor checks proposed actions against invariants and blocks violations. Makes no claim about the network's reasoning — only that certain outcomes cannot occur, which is the claim a safety case needs.
Prove that no input inside a stated region produces an output outside a stated one. A real proof about a real network, over a region you chose — which is the catch. It says nothing about inputs outside it, so the value of the guarantee is the value of the region, and that is a modelling judgement, not a solver result.
Katz et al., Reluplex: An Efficient SMT Solver for Verifying Deep Neural Networks
Turn any model's scores into a set of answers that contains the truth at a rate you choose, with no assumption about the model. The guarantee is distribution-free and holds on average over exchangeable data, which is a far weaker claim than it first sounds and far stronger than anything a bare softmax offers. When the model is unsure the set gets larger, which is the honest failure signal a single number cannot give.
Angelopoulos and Bates, A Gentle Introduction to Conformal Prediction
The structure here is ours: six parts and twenty-four chapters, arranged around what a reader has to be able to justify rather than around any one book. Each chapter names several works and courses that cover the same ground, so the reader can go to whichever one they already own. Nothing from any of them is reproduced; every unit is written for this site.
Courses and lecture series that cover a chapter. Every URL was fetched and its page title checked before it was listed; a chapter with nothing here had nothing we could verify, which is more useful to know than a plausible guess. Nothing is endorsed, affiliated or paid.
Trending papers, grouped by the week they were published. Selection is community signal, not our judgement, and each entry links to the source.
What this is and is not. Papers are harvested from Hugging Face Papers, ranked by that community's upvotes. Summaries are the authors' own abstracts, shortened — not our paraphrase, and not machine-generated commentary. The list reflects what is trending now rather than a persisted archive: older weeks appear only while they are still trending.