数组

数组 属于 Python

96 个练习

关于 数组

list是一个可变的集合,其中的元素按_顺序_排列。 和大多数集合一样(参见内置类型 tuple、dict 和 set),数组可以保存任意一种(或多种)数据类型的引用,包括其他数组。 和任何序列一样,元素可以从左侧通过0-based index下标访问,也可以从右侧通过-1-based index下标访问。 数组可以通过切片语法或<list>.copy()整体或部分复制。

数组同时支持通用序列操作和可变序列操作,例如min()/max()、<list>.index()、.append()和.reverse()。 可以使用for item in <list>结构遍历数组元素。当既需要元素下标又需要元素值时,可以使用for index, item in enumerate(<list>)。

数组的实现方式是动态数组,类似 Java 的Arraylist类型。它最常用于存储长度未知(条目数量可以任意增加或减少)的同类数据(字符串、数字、集合等)。

访问元素、通过in检查是否包含某个元素,或者在数组的“右侧”追加元素,都非常高效。 而在数组前面插入(即向“左侧”追加),或者插入到数组中间,效率则_低_得多,因为这些操作需要移动元素以保持它们的顺序。 如果想要一种类似的数据结构,能够从两端都高效地appends/pops,请参见collections.deque,它在两个方向上的性能都接近 O(1)。

由于数组是可变的,并且可以包含对任意 Python 对象的引用,因此在长度看似相同时,它们占用的内存也比array.array或tuple(它是不可变的)更多。 尽管如此,数组仍是一种极其灵活、有用的数据结构,Python 中许多内置方法和操作都会以数组作为输出结果。

构造

list可以通过_字面量_声明,使用方括号[],元素之间用逗号分隔:

>>> no_elements = []

>>> no_elements
[]

>>> one_element = ["Guava"]

>>> one_element
['Guava']

>>> elements_separated_with_commas = ["Parrot", "Bird", 334782]

>>> elements_separated_with_commas
['Parrot', 'Bird', 334782]

为了提高可读性,当数组中元素很多或存在嵌套数据结构时,可以使用换行:

>>> lots_of_entries = [
...    "Rose",
...    "Sunflower",
...    "Poppy",
...    "Pansy",
...    "Tulip",
...    "Fuchsia",
...    "Cyclamen",
...    "Lavender"
... ]

>>> lots_of_entries
['Rose', 'Sunflower', 'Poppy', 'Pansy', 'Tulip', 'Fuchsia', 'Cyclamen', 'Lavender']


# Each data structure is on its own line to help clarify what they are.
>>> nested_data_structures = [
...    {"fish": "gold", "monkey": "brown", "parrot": "grey"},
...    ("fish", "mammal", "bird"),
...    ['water', 'jungle', 'sky']
... ]

>>> nested_data_structures
[{'fish': 'gold', 'monkey': 'brown', 'parrot': 'grey'}, ('fish', 'mammal', 'bird'), ['water', 'jungle', 'sky']]

list()构造函数可以不带参数使用,也可以传入一个_可迭代对象_作为参数。 构造函数会依次遍历可迭代对象中的元素,并按顺序把它们加入数组:

>>> no_elements = list()
>>> no_elements
[]

# The tuple is unpacked and each element is added.
>>> multiple_elements_from_tuple = list(("Parrot", "Bird", 334782))

>>> multiple_elements_from_tuple
['Parrot', 'Bird', 334782]

# The set is unpacked and each element is added.
>>> multiple_elements_from_set = list({2, 3, 5, 7, 11})

>>> multiple_elements_from_set
[2, 3, 5, 7, 11]

用字符串或字典调用数组构造函数时,结果可能出乎意料:

# String elements (Unicode code points) are iterated through and added *individually*.
>>> multiple_elements_string = list("Timbuktu")

>>> multiple_elements_string
['T', 'i', 'm', 'b', 'u', 'k', 't', 'u']

# Unicode separators and positioning code points are also added *individually*.
>>> multiple_code_points_string = list('अभ्यास')

>>> multiple_code_points_string
['अ', 'भ', '्', 'य', 'ा', 'स']

# The iteration default for dictionaries is over the keys, so only key data is inserted into the list.
>>> source_data = {"fish": "gold", "monkey": "brown"}
>>> list(source_data)
['fish', 'monkey']

由于list()构造函数只接受可迭代对象(或者不传参数),不可迭代的对象会抛出TypeError。因此,用字面量方式创建只含一个元素的数组要容易得多。

# Numbers are not iterable, and so attempting to create a list with a number passed to the constructor fails.
>>> one_element = list(16)
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
TypeError: 'int' object is not iterable

# Tuples *are* iterable, so passing a one-element tuple to the constructor does work, but it's awkward
>>> one_element_from_iterable = list((16,))

>>> one_element_from_iterable
[16]

访问元素

数组中的元素(以及其他序列类型如str和tuple中的元素)可以使用_方括号表示法_访问。 下标可以从 left --> right(从 0 开始),也可以从 right --> left(从 -1 开始)。

从左数起的下标 ⟹






0
👇🏾
1
👇🏾
2
👇🏾
3
👇🏾
4
👇🏾
5
👇🏾
P y t h o n
👆🏾
-6
👆🏾
-5
👆🏾
-4
👆🏾
-3
👆🏾
-2
👆🏾
-1





⟸ 从右数起的下标
>>> breakfast_foods = ["Oatmeal", "Fruit Salad", "Eggs", "Toast"]

# Oatmeal is at index 0 or index -4.
>>> breakfast_foods[0]
'Oatmeal'

>>> breakfast_foods[-4]
'Oatmeal'

# Eggs are at index -2 or 2
>>> breakfast_foods[-2]
'Eggs'

>>> breakfast_foods[2]
'Eggs'

# Toast is at -1
>>> breakfast_foods[-1]
'Toast'

数组的一部分可以通过_切片语法_(<list>[<start>:<stop>])访问。 _切片_定义为位于index位置的一段元素序列,满足start <= index < stop。 切片返回被“切出”元素的副本,不会修改原来的list。

切片中还可以使用step参数(<list>[<start>:<stop>:<step>]),用来“跳过”或筛选返回的元素(例如,step为 2 时会选取该段中每隔一个的元素):

>>> colors = ["Red", "Purple", "Green", "Yellow", "Orange", "Pink", "Blue", "Grey"]

# If there is no step parameter, the step is assumed to be 1.
>>> middle_colors = colors[2:6]

>>> middle_colors
['Green', 'Yellow', 'Orange', 'Pink']

# If the start or stop parameters are omitted, the slice will
# start at index zero, and will stop at the end of the list.
>>> primary_colors = colors[::3]

>>> primary_colors
['Red', 'Yellow', 'Blue']

使用数组

数组提供了一个迭代器,可以像其他_序列类型_一样用for item in <list>或for index, item in enumerate(<list>)进行遍历:

# Make a list, and then loop through it to print out the elements
>>> colors = ["Orange", "Green", "Grey", "Blue"]
>>> for item in colors:
...     print(item)

Orange
Green
Grey
Blue


# Print the same list, but with the indexes of the colors included
>>> colors = ["Orange", "Green", "Grey", "Blue"]
>>> for index, item in enumerate(colors):
...     print(item, ":", index)

Orange : 0
Green : 1
Grey : 2
Blue : 3


# Start with a list of numbers and then loop through and print out their cubes.
>>> numbers_to_cube = [5, 13, 12, 16]
>>> for number in numbers_to_cube:
...     print(number**3)

125
2197
1728
4096

一种常见的构建数组的方式是在循环中使用<list>.append():

>>> cubes_to_1000 = []
>>> for number in range(11):
...    cubes_to_1000.append(number**3)

>>> cubes_to_1000
[0, 1, 8, 27, 64, 125, 216, 343, 512, 729, 1000]

数组还可以通过多种技巧合并:

# Using the plus + operator unpacks each list and creates a new list, but it is not efficient.
>>> new_via_concatenate = ["George", 5] + ["cat", "Tabby"]

>>> new_via_concatenate
['George', 5, 'cat', 'Tabby']

# Likewise, using the multiplication operator * is the equivalent of using + n times.
>>> first_group = ["cat", "dog", "elephant"]
>>> multiplied_group = first_group * 3

>>> multiplied_group
['cat', 'dog', 'elephant', 'cat', 'dog', 'elephant', 'cat', 'dog', 'elephant']

# Another method for combining 2 lists is to use slice assignment or a loop-append.
# This assigns the second list to index 0 in the first list.
>>> first_one = ["cat", "Tabby"]
>>> second_one = ["George", 5]
>>> first_one[0:0] = second_one

>>> first_one
['George', 5, 'cat', 'Tabby']

# This loops through the first list and appends its items to the end of the second list.
>>> first_one = ["cat", "Tabby"]
>>> second_one = ["George", 5]

>>> for item in first_one:
...      second_one.append(item)

>>> second_one
['George', 5, 'cat', 'Tabby']

一些注意事项

回忆一下,Python 中的变量是指向_底层对象_的_标签_。 lists多了一层身份,即_容器对象_:它们为收集到的元素保存对象的_引用_。 如果不正确处理,在使用数组时可能会引发多种潜在问题。

把同一个对象赋给多个变量名

把list对象赋给一个新的变量_名_,不会复制该list对象,也不会复制它的元素。 通过_新_变量名对list中元素所做的任何修改,都会_影响原来的对象_。

通过list.copy()或切片做一个shallow_copy,可以避免这一层引用带来的麻烦。 shallow_copy会创建一个新的list对象,但不会为其中包含的数组_元素_创建新对象。这种复制通常足以让你独立地向两个list对象添加或删除元素,从而实际上拥有两个“独立”的数组。

>>> actual_names = ["Tony", "Natasha", "Thor", "Bruce"]

# Assigning a new variable name does not make a copy of the container or its data.
>>> same_list = actual_names

#  Altering the list via the new name is the same as altering the list via the old name.
>>> same_list.append("Clarke")
["Tony", "Natasha", "Thor", "Bruce", "Clarke"]

>>> actual_names
["Tony", "Natasha", "Thor", "Bruce", "Clarke"]

#  Likewise, altering the data in the list via the original name will also alter the data under the new name.
>>> actual_names[0] = "Wanda"
['Wanda', 'Natasha', 'Thor', 'Bruce', 'Clarke']

# If you copy the list, there will be two separate list objects which can be changed independently.
>>> copied_list = actual_names.copy()
>>> copied_list[0] = "Tony"

>>> actual_names
['Wanda', 'Natasha', 'Thor', 'Bruce', 'Clarke']

>>> copied_list
["Tony", "Natasha", "Thor", "Bruce", "Clarke"]

这种引用带来的麻烦,在处理嵌套或重复相乘的数组时会变得更加严重(下面的例子来自 2013 年 Ned Batchelder 那篇精彩的博文 Names and values: making a game board):

from pprint import pprint

# This will produce a game grid that is 8x8, pre-populated with zeros.
>>> game_grid = [[0]*8]*8

>>> pprint(game_grid)
[[0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0]]

# An attempt to put a "X" in the bottom right corner.
>>> game_grid[7][7] = "X"

# This attempt doesn't work because all the rows are referencing the same underlying list object.
>>> pprint(game_grid)
[[0, 0, 0, 0, 0, 0, 0, 'X'],
 [0, 0, 0, 0, 0, 0, 0, 'X'],
 [0, 0, 0, 0, 0, 0, 0, 'X'],
 [0, 0, 0, 0, 0, 0, 0, 'X'],
 [0, 0, 0, 0, 0, 0, 0, 'X'],
 [0, 0, 0, 0, 0, 0, 0, 'X'],
 [0, 0, 0, 0, 0, 0, 0, 'X'],
 [0, 0, 0, 0, 0, 0, 0, 'X']]

不过在这种情况下,一个shallow_copy就足以实现我们想要的行为:

from pprint import pprint

# This loop will safely produce a game grid that is 8x8, pre-populated with zeros
>>> game_grid = []
>>> filled_row = [0] * 8
>>> for row in range(8):
...    game_grid.append(filled_row.copy()) # This is making a new shallow copy of the inner list object each iteration.

>>> pprint(game_grid)
[[0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0]]

# An attempt to put a "X" in the bottom right corner.
>>> game_grid[7][7] = "X"

# The game grid now works the way we expect it to!
>>> pprint(game_grid)
[[0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 'X']]

如前所述,数组是_引用_的容器,因此还有第二层潜在麻烦。 如果数组中包含变量、对象或嵌套数据结构,那么这些第二层引用不会通过shallow_copy或切片被复制。 修改底层对象会影响到_所有_副本,因为每个list对象只包含_指向_所存元素的_引用_。

from pprint import pprint

>>> pprint(game_grid)
[[0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 'X']]

# We'd like a new board, so we make a shallow copy.
>>> new_game_grid = game_grid.copy()

# But a shallow copy doesn't copy the contained references or objects.
>>> new_game_grid[0][0] = 'X'

# So changing the items in the copy also changes the originals items.
>>>  pprint(game_grid)
[['X', 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 0, 0, 0, 0, 0, 'X']]

相关数据类型

数组常被用作_栈_和_队列_,尽管其底层实现使得在前面插入和中间插入都比较慢。 collections模块提供了一个deque变体,它针对从两端快速追加和弹出做了优化,并以双向链表实现。 嵌套数组也常被用来表示小型_矩阵_,不过 Numpy 和 Pandas 库在进行高效的矩阵和表格数据处理方面要强大得多。 collections 模块还提供了UserList类型,可以根据特殊的数组需求进行定制。

通过 GitHub 编辑 该链接会在新窗口或标签页中打开

学习 数组

练习已锁定

再解锁 5 个练习即可练习 数组