如何阅读 Python 回溯信息以进行调试的概览。
调用函数时,会创建一个帧对象,用来保存局部变量和传给函数的实参。
函数返回时,帧对象会被销毁。
在函数A内调用函数B时,函数B的值会被放进一个帧对象,这个帧对象随后被放到调用栈中函数A的帧对象之上。
调用栈是当前处于活动状态的函数所对应的帧对象的集合。
如果函数A调用了函数B,而函数B又调用了函数C,那么这三个函数的帧对象都会出现在调用栈中。
函数C一旦返回,它的帧对象就会从栈中弹出,调用栈上只剩下函数A和B的帧对象。
回溯是某一时刻栈上所有帧对象的报告。 当 Python 程序遇到未处理的异常时,会打印异常信息和一份回溯。 回溯会显示异常在哪里被抛出,以及在此之前调用了哪些函数。
ValueError是一种常见的异常。
下面是一个ValueError的例子,它源于试图用右边的一个值给左边的两个变量赋值:
>>> first, second = [1]
Traceback (most recent call last):
File <stdin>, line 1, in <module>
first, second = [1]
ValueError: not enough values to unpack (expected 2, got 1)
回溯的排列方式是最近一次调用在最后,所以阅读回溯应该从最底部的异常开始。 从那里往上读,就能看出那条语句是怎样执行到的。 如果把出错的那一行放进一个函数里再调用这个函数,就会看到更长的回溯:
>>> def my_func():
... first, second = [1]
...
>>> my_func()
Traceback (most recent call last):
File <stdin>, line 5, in <module>
my_func()
File <stdin>, line 2, in my_func
first, second = [1]
ValueError: not enough values to unpack (expected 2, got 1)
从底部往回看,可以看到异常发生的那次调用位于my_func的第 2 行。
我们是通过在第 5 行调用my_func到达那里的。
Python 定义了 60 多个内置异常类。 下面简要介绍一些比较常见的异常,以及它们各自说明什么。
当 Python 因为语法无效而无法理解代码时,会抛出SyntaxError。
例如,可能有一个左括号没有与之匹配的右括号。
运行以下代码:
def distance(strand_a, strand_b):
if len(strand_a) != len(strand_b):
raise ValueError("Strands must be of equal length." # This is missing the closing parenthesis
会得到类似下面这样的堆栈跟踪(注意最后一行的信息):
.usr.local.lib.python3.10.site-packages._pytest.python.py:608: in _importtestmodule
mod = import_path(self.path, mode=importmode, root=self.config.rootpath)
.usr.local.lib.python3.10.site-packages._pytest.pathlib.py:533: in import_path
importlib.import_module(module_name)
.usr.local.lib.python3.10.importlib.__init__.py:126: in import_module
return _bootstrap._gcd_import(name[level:], package, level)
<frozen importlib._bootstrap>:1050: in _gcd_import ???
<frozen importlib._bootstrap>:1027: in _find_and_load ???
<frozen importlib._bootstrap>:1006: in _find_and_load_unlocked ???
<frozen importlib._bootstrap>:688: in _load_unlocked ???
.usr.local.lib.python3.10.site-packages._pytest.assertion.rewrite.py:168: in exec_module
exec(co, module.__dict__)
.mnt.exercism-iteration.hamming_test.py:3: in <module>
from hamming import (
E File ".mnt.exercism-iteration.hamming.py", line 10
E raise ValueError("Strands must be of equal length."
E ^
E SyntaxError: '(' was never closed
当assert语句(见下文)失败时,Python 会抛出AssertionError。
运行以下代码:
def distance(strand_a, strand_b):
assert len(strand_a) == len(strand_b)
distance("ab", "abc")
会得到类似下面这样的堆栈跟踪(注意最后一行的信息):
hamming_test.py:3: in <module>
from hamming import (
hamming.py:5: in <module>
distance("ab", "abc")
hamming.py:2: in distance
assert len(strand_a) == len(strand_b)
E AssertionError
当代码(或者单元测试!)试图访问某个对象的属性,而该对象没有这个属性时,就会抛出AttributeError。
例如,某个单元测试期望一个Robot对象有direction属性,但当它试图访问robot.direction时,这个属性并不存在。
这也可能说明有拼写错误,比如用了"Hello".lowercase(),而正确的写法是"Hello".lower()。
"Hello".lowercase()会抛出AttributeError: 'str' object has no attribute 'lowercase'。
运行以下代码:
class Robot:
def __init__():
#note that there is no self.direction listed here
self.position = (0, 0)
self.orientation = 'SW'
def forward():
pass
robby = Robot
robby.direction
会得到类似下面这样的堆栈跟踪(注意最后一行的信息):
robot_simulator_test.py:3: in <module>
from robot_simulator import (
robot_simulator.py:12: in <module>
robby.direction
E AttributeError: type object 'Robot' has no attribute 'direction'
运行以下代码:
def distance(strand_a, strand_b):
if strand_a.lowercase() == strand_b:
return 0
distance("ab", "abc")
会得到类似下面这样的堆栈跟踪(注意最后一行的信息):
def distance(strand_a, strand_b):
> if strand_a.lowercase() == strand_b:
E AttributeError: 'str' object has no attribute 'lowercase'
当代码试图导入某个东西,而 Python 无法做到时,就会抛出ImportError。
例如,Guidos Gorgeous Lasagna的单元测试会执行from lasagna import bake_time_remaining,但lasgana.py解答文件里可能没有定义bake_time_remaining。
运行lasgana.py文件但没有定义那个函数,会得到下面这样的错误:
We received the following error when we ran your code:
ImportError while importing test module '.mnt.exercism-iteration.lasagna_test.py'.
Hint: make sure your test modules.packages have valid Python names.
Traceback:
.mnt.exercism-iteration.lasagna_test.py:6: in <module>
from lasagna import (EXPECTED_BAKE_TIME,
E ImportError: cannot import name 'bake_time_remaining' from 'lasagna' (.mnt.exercism-iteration.lasagna.py)
During handling of the above exception, another exception occurred:
.usr.local.lib.python3.10.importlib.__init__.py:126: in import_module
return _bootstrap._gcd_import(name[level:], package, level)
.mnt.exercism-iteration.lasagna_test.py:23: in <module>
raise ImportError("In your 'lasagna.py' file, we can not find or import the"
E ImportError: In your 'lasagna.py' file, we can not find or import the function named 'bake_time_remaining()'. Did you mis-name or forget to define it?
### **IndexError**
Python raises an `IndexError` when an invalid index is used to look up a value in a sequence.
This often indicates the index is not computed properly and is often an off-by-one error.
<details>
<summary>Click here for code example</summary>
Consider the following code.
```python
def distance(strand_a, strand_b):
same = 0
for i in range(len(strand_a)):
if strand_a[i] == strand_b[i]:
same += 1
return same
distance("abc", "ab") # Note the first strand is longer than the second strand.
运行这段代码会得到类似下面这样的错误。 (注意最后一行。)
hamming_test.py:3: in <module>
from hamming import (
hamming.py:9: in <module>
distance("abc", "ab") # Note the first strand is longer than the second strand.
hamming.py:4: in distance
if strand_a[i] == strand_b[i]:
E IndexError: string index out of range
与IndexError类似,当用某个键去查找字典中的值,而这个键并没有设置在字典里时,就会抛出这个异常。
看看下面的代码。
def to_rna(dna_letter):
translation = {"G": "C", "C": "G", "A": "U", "T": "A"}
return translation[dna_letter]
print(to_rna("Q")) # Note, "Q" is not in the translation.
运行这段代码会得到类似下面这样的错误。 (注意最后一行。)
rna_transcription_test.py:3: in <module>
from rna_transcription import to_rna
rna_transcription.py:6: in <module>
print(to_rna("Q"))
rna_transcription.py:3: in to_rna
return translation[dna_letter]
E KeyError: 'Q'
TypeError通常在传给函数或在运算中使用的数据类型不对时抛出。
看看下面的代码。
def hello(name): # This function expects a string.
return 'Hello, ' + name + '!'
print(hello(100)) # 100 is not a string.
运行这段代码会得到类似下面这样的错误。 (注意最后一行。)
hello_world_test.py:3: in <module>
import hello_world
hello_world.py:5: in <module>
print(hello(100))
hello_world.py:2: in hello
return 'Hello, ' + name + '!'
E TypeError: can only concatenate str (not "int") to str
当把无效的值传给函数时,通常会抛出ValueError。
注意,实数平方根只对正数存在。
调用math.sqrt(-1)会抛出ValueError: math domain error,因为-1不是平方根的有效取值。
用(数学上的)术语来说,-1 不在平方根的定义域内。
import math
math.sqrt(-1)
运行这段代码会得到类似下面这样的错误。 (注意最后一行。)
square_root_test.py:3: in <module>
from square_root import (
square_root.py:3: in <module>
math.sqrt(-1)
E ValueError: math domain error
print函数有时候并没有抛出错误,但某个值却不是预期的结果。 如果这个值是经过一连串计算得出的,就会格外令人困惑。 在这种情况下,查看每一步的值会很有帮助,能看出是哪一步没有按预期运行。 可以用 print 函数把值打印到控制台。 下面的例子是一个没有返回预期值的函数:
# the intent is to pass an integer to this function and get an integer back
def halve_and_quadruple(num):
return (num / 2) * 4
把5传给这个函数时,预期值是8,但它返回的是10.0。
为了排查,把计算拆解开,这样就可以在每一步检查这个值。
# the intent is to pass an integer to this function and get an integer back
def halve_and_quadruple(num):
# verify the number in is what is expected
# prints 5
print(num)
# we want the int divided by an integer to be an integer
# but this prints 2.5! We've found our mistake.
print(num / 2)
# this makes sense, since 2.5 x 4 = 10.0
print((num / 2) * 4)
return (num / 2) * 4
What the `print` calls revealed is that we used `/` when we should have used `//`, the [floor division operator][floor division operator].
## Logging
[Logging][logging] can be used similarly to `print`, but it is more powerful.
What is logged can be configured by the logging severity (e.g., 'DEBUG', 'INFO', 'WARNING', 'ERROR', 'CRITICAL'.)
A call to the `logging.error` function can pass `True` to the `exc_info` parameter, which will additionally log the stack trace.
By configuring multiple handlers, logging can write to more than one place with the same logging function.
Following is an example of logging printed to the console:
```python
>>> import logging
>>>
>>> # configures minimum logging level as INFO
>>> logging.basicConfig(level=logging.INFO)
>>>
>>> def halve_and_quadruple(num):
... # prints INFO:root: num == 5
... logging.info(f" num == {num}")
... return (num // 2) * 4
...
>>> print(halve_and_quadruple(5))
级别配置为INFO,因为默认级别是WARNING。
如果要让日志持久保存,可以把记录器配置为写入文件,像这样:
>>> import logging
...
>>> # configures the output file name to example.log, and the minimum logging level as INFO
>>> logging.basicConfig(filename='example.log', level=logging.INFO)
...
... def halve_and_quadruple(num):
... # prints INFO:root: num == 5 to the example.log file
... logging.info(f" num == {num}")
... return (num // 2) * 4
...
>>> print(halve_and_quadruple(5))
assert 是一条语句,除非程序里有 bug,否则它应该始终求值为True。
当assert求值为False时,它会抛出AssertionError。
AssertionError的回溯可以包含一条可选的信息,它是assert语句的一部分。
虽然信息是可选的,但最好总是在assert定义里加上一条。
下面是一个使用assert的例子:
>>> def int_division(dividend, divisor):
... assert divisor != 0, "divisor must not be 0"
... return dividend // divisor
...
>>> print(int_division(2, 1))
2
>>> print(int_division(2, 0))
Traceback (most recent call last):
File <stdin>, line 7, in <module>
print(int_division(2, 0))
^^^^^^^^^^^^^^^^^^
File <stdin>, line 2, in int_division
assert divisor != 0, "divisor must not be 0"
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
AssertionError: divisor must not be 0
如果我们像应该做的那样从底部开始阅读回溯,很快就能看出问题在于不该把0作为divisor传进去。
assert也可以用来检查一个值是否属于预期的类型:
>>> import numbers
...
...
... def int_division(dividend, divisor):
... assert divisor != 0, "divisor must not be 0"
... assert isinstance(divisor, numbers.Number), "divisor must be a number"
... return dividend // divisor
...
>>> print(int_division(2, 1))
2
>>> print(int_division(2, '0'))
Traceback (most recent call last):
File <stdin>, line 11, in <module>
print(int_division(2, '0'))
^^^^^^^^^^^^^^^^^^^^
File <stdin>, line 6, in int_division
assert isinstance(divisor, numbers.Number), "divisor must be a number"
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
AssertionError: divisor must be a number
一旦找出 bug,可以考虑用错误处理来替代assert。
这是因为可以通过用-O或-OO选项运行 Python,或者把PYTHONOPTIMIZE环境变量设为1或2,来禁用所有的assert语句。
把PYTHONOPTIMIZE设为1等同于用-O选项运行 Python,它会禁用断言。
把PYTHONOPTIMIZE设为2等同于用-OO选项运行 Python,它既禁用断言,又会从字节码中移除文档字符串。
减少字节码是让代码运行得更快的一种方法。
Python 有一个内置调试器 pdb。
可以用它单步执行代码并检查变量。
还可以用它设置断点。
开始时,你需要先import pdb,然后在想要开始调试的地方调用pdb.set_trace():
import pdb
def add(num1, num2):
return num1 + num2
pdb.set_trace()
sum = add(1,5)
print(sum)
运行这段代码会得到一个 pdb 提示符,可以在里面输入命令。
输入help可以获取命令列表。
最常用的命令是step,它会进入该行所调用的函数内部。
next会跳过函数调用,移到下一行。where会告诉你当前位于哪一行。
其他一些有用的命令有whatis <variable>,它会告诉你变量的类型,以及print(<variable>),它会打印变量的值。
你也可以直接用<variable>打印变量的值。
另一个命令是jump <line number>,它会跳转到指定的行号。
下面是一个基于前面代码的小例子,演示如何使用调试器。 注意,在这个以及后面的例子里,MacOS 或 Linux 平台上的文件路径会使用正斜杠:
>>> python pdb.py
... > c:\pdb.py(7)<module>()
... -> sum = add(1,5)
... (Pdb)
>>> step
... > c:\pdb.py(3)add()
... -> def add(num1, num2):
... (Pdb)
>>> whatis num1
... <class 'int'>
>>> print(num2)
... 5
>>> next
... > c:\pdb.py(4)add()
... -> return num1 + num2
... (Pdb)
>>> jump 3
... > c:\pdb.py(3)add()
... -> def add(num1, num2):
... (Pdb)
断点通过break <filename>:<line number> <condition>来设置,其中 condition 是一个可选条件,只有它为真时断点才会被命中。
直接输入break就能列出你已经设置的断点。
要禁用某个断点,可以输入disable <breakpoint number>。
要启用某个断点,可以输入enable <breakpoint number>。
要删除某个断点,可以输入clear <breakpoint number>。
要继续执行,可以输入continue或c。要退出调试器,可以输入quit或q。
下面是一个基于前面代码的例子,演示如何使用上述调试器命令:
>>> python pdb.py
... > c:\pdb.py(7)<module>()
... -> sum = add(1,5)
... (Pdb)
>>> break
...
>>> break pdb:4
... Breakpoint 1 at c:\pdb.py:4
>>> break
... Num Type Disp Enb Where
... 2 breakpoint keep yes at c:\pdb.py:4
>>> c # continue
... > c:\pdn.py(4)add()
... -> return num1 + num2
>>> disable break 1
... Disabled breakpoint 1 at c:\pdb.py:4
>>> break
... Num Type Disp Enb Where
... 1 breakpoint keep no at c:\pdb.py:4
... breakpoint already hit 1 time
>>> clear break 1
... Deleted breakpoint 1 at c:\pdb.py:4
>>> break
...
在 Python 3.7+ 中,有一种更简单的创建断点的方式。
只需在需要的地方写上breakpoint()就能创建一个断点。
def add(num1, num2):
breakpoint()
return num1 + num2
breakpoint()
sum = add(1,5)
print(sum)
>>> python pdb.py
... > c:\pdb.py(7)<module>()
... -> sum = add(1,5)
... (Pdb)
>>> c # continue
... > c:\pdb.py(5)add()
... -> return num1 + num2