Machine learning types and training · 机器学习类型与训练
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| machine learning/məˈʃiːn ˈlɜːnɪŋ/ | 机器学习 | jī qì xué xí |
| backpropagation/ˌbækprəpəˈɡeɪʃn/ | 反向传播 | fǎn xiàng chuán bō |
| supervised learning/ˈsuːpəvaɪzd ˈlɜːnɪŋ/ | 监督学习 | jiān dū xué xí |
| classification/ˌklæsɪfɪˈkeɪʃn/ | 分类 | fēn lèi |
| unsupervised learning/ʌnˈsuːpəvaɪzd ˈlɜːnɪŋ/ | 无监督学习 | wú jiān dū xué xí |
| clusters/ˈklʌstəz/ | 聚类 | jù lèi |
| reinforcement learning/ˌriːɪnˈfɔːsmənt ˈlɜːnɪŋ/ | 强化学习 | qiáng huà xué xí |
| agent/ˈeɪdʒənt/ | 智能体 | zhì néng tǐ |
| reward/rɪˈwɔːd/ | 奖励 | jiǎng lì |
| policy/ˈpɒlɪsi/ | 策略 | cè lüè |
| gradient descent/ˈɡreɪdɪənt dɪˈsent/ | 梯度下降 | tī dù xià jiàng |
| loss function/lɒs ˈfʌŋkʃn/ | 损失函数 | sǔn shī hán shù |
| epoch/ˈiːpɒk/ | 训练轮次 | xùn liàn lún cì |
The program that taught itself an opening no human had played
- AlphaGo Zero was given the rules of Go and nothing else. No human games, no opening book, no advice. It played against itself, starting from random moves.
- In three days it surpassed the version that had beaten the world champion. Along the way it rediscovered centuries of human opening theory, and then discarded parts of it as inferior.
- It learned from a single number after each game: won or lost. That is a completely different kind of learning from being shown a million labelled photographs.
- This lesson is the three paradigms of machine learning, what data each needs, and how a network's weights are actually changed by backpropagation 反向传播.
自学出人类没下过的定式的程序
- AlphaGo Zero 被给予的只有围棋规则,别无其他。没有人类棋谱,没有定式书,没有指点。它从随机落子开始,自己与自己对弈。
- 三天之内,它超过了那个曾击败世界冠军的版本。一路上它重新发现了人类几百年的开局理论,然后又把其中一部分当作次优抛弃了。
- 它每盘棋只从一个数字里学习:赢了还是输了。这与被看一百万张有标签的照片是完全不同的一种学习。
- 这一课讲机器学习的三种范式、各需要什么数据,以及网络的权重究竟怎样被反向传播(backpropagation)改变。
Supervised learning
- Supervised learning 监督学习 trains on data that carries labels: photographs tagged "cat" or "dog", emails marked spam or not, houses with their sale prices.
- The model learns the mapping from input to label, and is then given inputs it has never seen and asked to produce the label.
- It splits into classification 分类, where the output is a category, and regression, where the output is a number.
- Its cost is the labels: someone must produce them, which for a million images is a great deal of human work.
Labelled examples in, a model out, then unlabelled inputs
监督学习
- 监督学习(supervised learning)在带标签的数据上训练:标注为"猫"或"狗"的照片、标为垃圾或非垃圾的邮件、带成交价的房屋。
- 模型学习从输入到标签的映射,然后被给予从未见过的输入,要求给出标签。
- 它分为输出是类别的分类(classification),和输出是数字的回归。
- 它的代价是标签:必须有人做出这些标签,而对一百万张图片来说那是巨大的人力。

带标签的样例进去,模型出来,再送入没有标签的输入
Supervised learning uses: · 监督学习使用:
Supervised learning trains on labelled examples (e.g. images tagged cat/dog). · 监督学习在标记示例上训练(例如标记为猫/狗的图片)。
Unsupervised learning
- Unsupervised learning 无监督学习 is given data with no labels and asked to find structure in it.
- The commonest use is clustering: grouping into clusters 聚类 of similar items, such as customers who buy similar things, without anyone deciding in advance what the groups should be.
- That is its strength and its weakness: it can discover a grouping nobody suspected, and it cannot tell you what the grouping means.
无监督学习
- 无监督学习(unsupervised learning)被给予没有标签的数据,要求在其中找出结构。
- 最常见的用途是聚类:把相似的项归成簇(clusters),比如买相似东西的顾客,而事先没有人规定这些组该是什么。
- 这既是它的长处也是它的短处:它能发现没人料到的分组,而它说不出这个分组意味着什么。
Unsupervised learning is used to: · 无监督学习用于:
With no labels, unsupervised learning discovers patterns/clusters in the data. · 在没有标签的情况下,无监督学习发现数据中的模式/聚类。
Reinforcement learning
- Reinforcement learning 强化学习 has an agent 智能体 acting in an environment. Each action changes the state and returns a reward 奖励, which may be zero for a long time.
- The agent learns a policy 策略, a strategy for choosing actions, that maximises its total reward over time. It learns by trial and error, with no labelled examples at all.
- It suits problems that are a sequence of decisions where the right move is only revealed by the eventual outcome: games, robot control, self-driving cars. That is exactly the AlphaGo Zero case.
强化学习
- 强化学习(reinforcement learning)里有一个在环境中行动的智能体(agent)。每个动作改变状态并返回一个奖励(reward),这个奖励可能很长时间都是零。
- 智能体学习一个策略(policy),即选择动作的方针,使它的长期总奖励最大。它靠试错学习,完全没有带标签的样例。
- 它适合一连串决策的问题,其中哪一步对只有最终结果才揭晓:下棋、机器人控制、自动驾驶。AlphaGo Zero 正是这种情形。
In reinforcement learning, the agent learns by: · 在强化学习中,智能体通过以下方式学习:
The agent tries actions, receives rewards, and learns a policy that maximises long-term reward. · 智能体尝试动作,接收奖励,并学习一个能最大化长期奖励的策略。
In reinforcement learning the agent learns a ____ that maximises its total reward over time. · 在强化学习中,智能体学习一个____以随时间推移最大化其总奖励。
It learns by trial and error from rewards, with no labelled examples at all, which is why it suits sequences of decisions. · 它通过奖励进行试错学习,完全不需要标记示例,因此适用于决策序列。
Worked example: choose the paradigm
- Ten thousand medical scans, each labelled by a doctor as showing a tumour or not, are used to train a system to flag new scans. Supervised: the data has labels and the task is to learn input to label, specifically a classification.
- A shop wants to know whether its customers fall into natural groups, without deciding the groups in advance. Unsupervised: no labels exist, and the task is to find structure, specifically clustering.
- A robot arm must learn to place components, judged only by whether each attempt succeeded. Reinforcement: an agent, a sequence of actions, and a reward rather than a label.
- Name the paradigm, then justify from the data: labelled, unlabelled, or a reward signal.
例题:选择范式
- 一万张医学扫描图,每张都由医生标注为有无肿瘤,用来训练一个系统对新扫描图作标记。 监督:数据有标签,任务是学习输入到标签的映射,具体说是分类。
- 一家商店想知道顾客是否落入若干自然的组,而不事先规定这些组。 无监督:没有标签,任务是找出结构,具体说是聚类。
- 一只机械臂必须学会放置零件,判断依据只有每次尝试是否成功。 强化:一个智能体、一连串动作,以及奖励而不是标签。
- 说出范式,再从数据来论证:有标签、无标签,还是一个奖励信号。
AI learning type lab · AI学习类型实验
Classify AI examples by the type of learning or concern involved. · 根据涉及的学习类型或关注点对AI示例进行分类。
Match each ML paradigm to its data. · 将每种机器学习范式与其所需的数据匹配。
Supervised = labels; unsupervised = no labels; reinforcement = reward feedback. · 监督学习=标签;无监督学习=无标签;强化学习=奖励反馈。
Match each situation to the machine-learning paradigm it needs. · 将每种情况与其所需的机器学习范式匹配。
Labels, no labels, or a reward. The data decides the paradigm, not the difficulty of the task. · 标签、无标签或奖励。数据决定范式,而非任务的难度。
Training by backpropagation
- Training means adjusting the weights so the network's outputs match the targets. The method is backpropagation with gradient descent 梯度下降.
- Forward pass: feed an input through the network and get its output. Compute the error using a loss function 损失函数, which measures how far the output is from the target.
- Backward pass: propagate that error backwards through the layers to find each weight's gradient, that is, how much that weight contributed to the error.
- Update: change each weight by a small step against its gradient. The step size is the learning rate.
Downhill, one small step at a time
用反向传播训练
- 训练就是调整权重,让网络的输出与目标相符。方法是配合梯度下降(gradient descent)的反向传播。
- 前向传播:把输入送过网络得到它的输出。用损失函数(loss function)计算误差,它衡量输出离目标有多远。
- 反向传播:把这个误差沿各层反向传播,求出每个权重的梯度,也就是那个权重对误差贡献了多少。
- 更新:让每个权重沿梯度的反方向走一小步。步长就是学习率。

向下走,一次一小步
Backpropagation trains a network by: · 反向传播通过以下方式训练网络:
The error flows from the output back through the network (chain rule) so every weight's gradient is found, then weights step downhill. · 误差从输出端沿网络反向流动(链式法则),从而找到每个权重的梯度,随后权重向下移动。
Put one round of backpropagation training in order. · 按顺序排列一轮反向传播训练过程。
Forward, error, backward, update, repeated over many epochs until the error stops falling. · 前向传递、误差计算、反向传递、更新,重复多个周期(epochs),直到误差不再下降。
Epochs, and why the learning rate matters
- One pass through the whole training set is an epoch 训练轮次. Training repeats for many epochs until the error stops falling.
- Too large a learning rate overshoots the minimum and the error jumps about; too small and training takes far longer than it needs to.
- After training, using the model is only a forward pass, which is why a trained network answers instantly even though training took days.
训练轮次,以及学习率为什么要紧
- 把整个训练集过一遍是一个训练轮次(epoch)。训练要重复很多轮,直到误差不再下降。
- 学习率太大会越过最小值,误差来回跳;太小则训练所需时间远超必要。
- 训练完成后,使用模型只需一次前向传播,这就是训练花了几天而训练好的网络能瞬间作答的原因。
In gradient descent the learning rate sets how big a step each weight update takes — too large overshoots the minimum, too small makes training slow. · 在梯度下降中,学习率设定了每次权重更新的步幅——过大则会越过最小值,过小则导致训练缓慢。
Backpropagation finds each weight's gradient; the learning rate scales how far down that gradient the weight moves. · 反向传播找到每个权重的梯度;学习率缩放权重沿该梯度移动的幅度。
Which statements about training are correct? Select all · 所有 that apply. · 关于训练哪些陈述是正确的?选择所有适用项。
The learning rate is the size of each weight update. That is why training can take days while a prediction takes milliseconds. · 学习率是每个权重更新的大小。这就是为什么训练可能需要数天,而预测仅需毫秒的原因。
Marks that slip away
- Justify a paradigm from the data: labelled means supervised, unlabelled means unsupervised, a reward signal means reinforcement.
- Reinforcement learning has no labels. Calling its reward a label is the classic confusion.
- Backpropagation is forward, error, backward, update, repeated. Naming only "it adjusts the weights" is half an answer.
- The learning rate is the size of each step, not the speed of training or the number of epochs.
容易丢掉的分
- 从数据论证范式:有标签就是监督,无标签就是无监督,有奖励信号就是强化。
- 强化学习没有标签。把它的奖励叫做标签是典型的混淆。
- 反向传播是前向、误差、反向、更新,不断重复。只说"它调整权重"只是半个答案。
- 学习率是每一步的大小,不是训练的速度,也不是轮次的数量。
You've got it
- supervised learns input to label from labelled data, for classification or regression · unsupervised finds structure such as clusters in unlabelled data · reinforcement has an agent learning a policy from rewards by trial and error
- choose the paradigm from what the data provides, not from the task's difficulty
- backpropagation: a forward pass, an error from the loss function, a backward pass giving each weight's gradient, then an update of size set by the learning rate
- training repeats for many epochs; using the trained model is one forward pass
你掌握了
- 监督从有标签数据学习输入到标签,用于分类或回归 · 无监督在无标签数据中找出结构比如簇 · 强化让智能体靠奖励试错学出一个策略
- 按数据提供了什么来选择范式,而不是按任务的难度
- 反向传播:一次前向传播、由损失函数得出的误差、一次给出每个权重梯度的反向传播,再按学习率决定的步长更新
- 训练重复很多轮次;使用训练好的模型只需一次前向传播