Compilation and interpretation · 编译与解释
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| compiler/kəmˈpaɪlə/ | 编译器 | biān yì qì |
| interpreter/ɪnˈtɜːprɪtə/ | 解释器 | jiě shì qì |
| lexical analysis/ˈleksɪkl əˈnæləsɪs/ | 词法分析 | cí fǎ fēn xī |
| tokens/ˈtəʊkənz/ | 词法单元 | cí fǎ dān yuán |
| syntax analysis/ˈsɪntæks əˈnæləsɪs/ | 语法分析 | yǔ fǎ fēn xī |
| abstract syntax tree/ˈæbstrækt ˈsɪntæks triː/ | 抽象语法树 | chōu xiàng yǔ fǎ shù |
| syntax error/ˈsɪntæks ˈerə/ | 语法错误 | yǔ fǎ cuò wù |
| semantic analysis/səˈmæntɪk əˈnæləsɪs/ | 语义分析 | yǔ yì fēn xī |
| code generation/kəʊd ˌdʒenəˈreɪʃn/ | 代码生成 | dài mǎ shēng chéng |
| code optimisation/kəʊd ˌɒptɪmaɪˈzeɪʃn/ | 代码优化 | dài mǎ yōu huà |
The bug that cost a spacecraft, found by a comma
- NASA's Mariner 1 was destroyed 293 seconds after launch in 1962. The failure is often blamed on a single wrong character in the guidance software.
- The story is retold so often because it is the fear every programmer knows: a compiler happily translating a program that says something you did not mean.
- A compiler is not one process but five, and each one catches a different class of mistake. Knowing which stage catches what tells you where any given error comes from.
- This lesson is how an interpreter 解释器 runs a program without producing anything, and the five stages a compiler 编译器 goes through.
据说被一个逗号毁掉的航天器
- NASA 的水手 1 号在 1962 年发射后 293 秒被摧毁。故障常被归咎于制导软件中的一个错字符。
- 这个故事被反复讲述,是因为它说中了每个程序员都懂的恐惧:编译器欣然翻译了一个说着你并不想说的话的程序。
- 编译器不是一个过程,而是五个,每一个抓住不同类别的错误。知道哪一阶段抓什么,就知道任何一个错误来自哪里。
- 这一课讲解释器(interpreter)怎样不产生任何东西就运行程序,以及编译器(compiler)经历的五个阶段。
How an interpreter runs a program
- An interpreter translates and runs the source at the same time, one statement at a time.
- For each statement it reads the line, analyses it, checks the types, then executes the action immediately, and moves to the next.
- It produces no executable file: the translation exists only in memory and is thrown away. That is the syllabus's exact point, that it can execute a program without producing a translated version.
- It reports an error the moment it reaches the offending line and usually stops, which gives quick feedback while developing. Every run re-translates, so it is slower, and both the source and the interpreter must be present.
Translate once and keep it, or translate and run and forget
解释器怎样运行程序
- 解释器同时翻译和运行源代码,一次一条语句。
- 对每条语句,它读入这一行、分析它、检查类型,然后立即执行这个动作,再转到下一条。
- 它不产生可执行文件:翻译只存在于内存中,随后被丢弃。这正是大纲的原话所指:它能执行程序而不产生翻译后的版本。
- 它一走到出错的那一行就报告错误并通常停下,这在开发时给出快速反馈。每次运行都重新翻译,所以更慢,而且源代码和解释器每次都必须在场。

翻译一次并保留,还是边翻译边运行再忘掉
An interpreter: · 一个解释器:
It works line by line, reporting errors as it reaches them; nothing is saved as an executable, and it is generally slower. · 它逐行工作,在到达错误时报告它们;没有东西被保存为可执行文件,而且它通常更慢。
An interpreter can execute a program without producing a translated version of it. · 解释器能执行一个程序而不产生它的翻译版本。
The translation exists only in memory, statement by statement, and is discarded. That is why the interpreter must be present every time the program runs. · 翻译只逐句存在于内存中,随后被丢弃。这就是每次运行程序时解释器都必须在场的原因。
Stage 1: lexical analysis
- Lexical analysis 词法分析 groups the individual characters of the source into tokens 词法单元: keywords, identifiers, operators and literals.
- Whitespace and comments are discarded, since they have no meaning to the compiler, and identifiers are entered into a symbol table.
- So
total ← count * 2becomes the token sequence identifier, assignment, identifier, operator, literal. The lexer does not know or care whether that makes sense.
阶段一:词法分析
- 词法分析(lexical analysis)把源代码中一个个字符归成词法单元(tokens):关键字、标识符、运算符和字面量。
- 空白和注释被丢弃,因为它们对编译器没有意义;标识符被登记进符号表。
- 于是
total ← count * 2变成词法单元序列:标识符、赋值号、标识符、运算符、字面量。词法分析器不知道也不关心这是否讲得通。
Lexical analysis (the lexer) turns: · 词法分析(词法分析器)把:
The lexer groups characters into tokens and discards whitespace/comments; parsing then builds the tree. · 词法分析器把字符分组成记号并丢弃空白/注释;解析然后构建树。
Lexical analysis groups the source characters into ____ and discards whitespace and comments. · 词法分析把源代码字符归成____,并丢弃空白和注释。
Keywords, identifiers, operators and literals. The lexer makes no judgement about whether the sequence is a valid program. · 关键字、标识符、运算符和字面量。词法分析器不判断这个序列是否是合法程序。
Stage 2: syntax analysis
- Syntax analysis 语法分析, or parsing, checks that the sequence of tokens fits the language's grammar, and builds an abstract syntax tree 抽象语法树 representing the structure.
- A missing bracket, a missing
ENDIFor a keyword in the wrong place is caught here: that is what a syntax error 语法错误 is, and it is why the compiler can report one without ever running the program.
The grammar the parser checks against
阶段二:语法分析
- 语法分析(syntax analysis,或称解析)检查词法单元序列是否符合语言的文法,并构建表示结构的抽象语法树(abstract syntax tree)。
- 缺一个括号、缺一个
ENDIF、关键字放错位置,都在这里被抓住:这就是语法错误(syntax error),也是编译器无需运行程序就能报错的原因。

解析器据以检查的文法
Stage 3: semantic analysis
- Semantic analysis 语义分析 checks that a syntactically valid program actually makes sense: every variable is declared before use, the types on each side of an assignment are compatible, a function is called with the right number of arguments.
total ← "seven" * 2is perfectly good syntax and complete nonsense. Only this stage catches it.- This is the distinction the exam tests: syntax is form, semantics is meaning.
阶段三:语义分析
- 语义分析(semantic analysis)检查一个语法上合法的程序是否真的讲得通:每个变量在使用前已声明、赋值两边的类型相容、函数调用的参数个数正确。
total ← "seven" * 2语法完美,语义完全说不通。只有这一阶段能抓住它。- 这正是考试考的区分:语法是形式,语义是含义。
Which of these is checked by semantic analysis rather than syntax analysis? · 以下哪一项由语义分析而不是语法分析检查?
Brackets are grammar, so syntax; whitespace is the lexer; registers are code generation. Declarations and types are meaning. · 括号是文法,属于语法;空白归词法分析器;寄存器属于代码生成。声明和类型属于含义。
Stages 4 and 5: code generation and optimisation
- Code generation 代码生成 walks the tree and emits the target machine code, choosing registers and laying out the data.
- Code optimisation 代码优化 improves that code without changing what it does: removing redundant work, computing constant expressions at compile time, and reordering instructions to suit the pipeline.
- The output is a stand-alone executable that runs without the compiler present.
Five stages, and each one catches something the last could not
阶段四和五:代码生成与优化
- 代码生成(code generation)遍历语法树并输出目标机器码,选择寄存器并安排数据布局。
- 代码优化(code optimisation)在不改变程序行为的前提下改进这些代码:去掉冗余工作、在编译时算出常量表达式、为流水线重排指令。
- 输出是一个独立的可执行文件,不需要编译器在场就能运行。

五个阶段,每一个抓住上一个抓不到的东西
The phases of compilation · 编译的各阶段
Step through what a compiler does to your source. Each phase hands its output to the next — characters become tokens, tokens become a tree, the tree becomes optimised machine code. · 逐步走过一个编译器对你的源代码做什么。每个阶段把它的输出交给下一个——字符变成记号,记号变成一棵树,树变成优化的机器码。
Match each compiler phase to what it does. · 把每个编译器阶段与它做的事配对。
Each phase transforms the program a step further: tokens, then a tree, then checked, then optimised code. · 每个阶段把程序进一步转换:记号,然后一棵树,然后被检查,然后优化的代码。
Put the compiler stages in order. · 把编译器阶段按顺序排列。
Lexical → syntax → semantic → code generation → optimisation. · 词法 → 语法 → 语义 → 代码生成 → 优化。
Code optimisation aims to: · 代码优化的目标是:
Optimisation improves the generated code (constant folding, removing redundancy, reordering for the pipeline). · 优化改进生成的代码(常量折叠、移除冗余、为流水线重新排序)。
Worked example: which stage catches which error
x ← 5 +: the tokens do not fit the grammar, an operator with no right operand, so syntax analysis.x ← y + 1whereywas never declared: the form is fine but the meaning is not, so semantic analysis.IF a > b THEN OUTPUT awith noENDIF: syntax analysis again.- A program that runs and prints the wrong average: no stage catches it. That is a logic error, and only testing finds it. Name the stage and why the earlier ones let it through.
例题:哪一阶段抓哪种错误
x ← 5 +:词法单元不符合文法,一个运算符没有右操作数,所以是语法分析。x ← y + 1而y从未声明:形式没问题,含义有问题,所以是语义分析。IF a > b THEN OUTPUT a没有ENDIF:又是语法分析。- 一个能运行却打印错误平均值的程序:没有任何阶段能抓住它。那是逻辑错误,只有测试才能发现。要说出阶段并且说出为什么前面的阶段放过了它。
Match each faulty line to the compiler stage that catches it. · 把每一行错误代码与抓住它的编译阶段配对。
Syntax is form, semantics is meaning, and a logic error is valid in both but wrong in intent. · 语法是形式,语义是含义,而逻辑错误在两方面都合法,只是意图错了。
Worked example: compare the two translators
- Give two differences between how a compiler and an interpreter handle a program. [4]
- A compiler translates the whole program before it runs and produces an executable, which then runs without the compiler; an interpreter translates and executes one statement at a time and produces no executable, so the interpreter must be present every run.
- A compiler reports all the errors together after translation; an interpreter reports the first error when it reaches that line and then stops.
- Each mark is a pairing: say what one does and what the other does instead.
例题:比较两种翻译器
- 给出编译器和解释器处理程序的两点区别。[4]
- 编译器在运行前翻译整个程序并产生一个可执行文件,之后运行时不需要编译器;解释器一次翻译并执行一条语句,不产生可执行文件,所以每次运行解释器都必须在场。
- 编译器在翻译结束后把错误全部一起报告;解释器在走到出错行时报告第一个错误然后停下。
- 每一分都是一组对照:说出一个怎么做,以及另一个改为怎么做。
Marks that slip away
- An interpreter produces no translated version. That is the phrase the syllabus uses and the mark it awards.
- The lexer produces tokens and discards whitespace and comments; it does not check whether the program is valid.
- Syntax is form, semantics is meaning. An undeclared variable is a semantic error, not a syntax error.
- Optimisation must not change what the program does, only how quickly or compactly it does it.
容易丢掉的分
- 解释器不产生翻译后的版本。这是大纲用的措辞,也是它给分的地方。
- 词法分析器产生词法单元并丢弃空白和注释;它不检查程序是否合法。
- 语法是形式,语义是含义。未声明的变量是语义错误,不是语法错误。
- 优化不得改变程序做什么,只能改变它做得多快、多紧凑。
You've got it
- an interpreter translates and executes one statement at a time, producing no executable, reporting the first error at its line, and re-translating on every run
- lexical analysis makes tokens and discards whitespace and comments; syntax analysis checks the grammar and builds an abstract syntax tree, catching syntax errors
- semantic analysis checks meaning: declarations, types, argument counts
- code generation emits machine code and optimisation improves it without changing behaviour, giving a stand-alone executable
你掌握了
- 解释器一次翻译并执行一条语句,不产生可执行文件,在出错行报告第一个错误,并且每次运行都重新翻译
- 词法分析产生词法单元并丢弃空白和注释;语法分析检查文法并构建抽象语法树,抓住语法错误
- 语义分析检查含义:声明、类型、参数个数
- 代码生成输出机器码,优化在不改变行为的前提下改进它,得到独立的可执行文件