字符

字符串 属于 Python

83 个练习

关于 字符串

Python 中的str是由 Unicode 码点组成的不可变序列。这些码点可以包括字母、变音符号、定位字符、数字、货币符号、emoji、标点、空格和换行符,等等。

想深入了解字符串到底编码了什么信息(或者说,“计算机怎么知道如何把 0 和 1 转换成字母?”),这篇博客文章始终很有帮助。Python 文档还提供了一份非常详细的 Unicode HOWTO,讨论了 Python 在str、bytes和re模块中对 Unicode 规范的支持、与区域设置相关的注意事项,以及编码和转换方面的一些常见问题。

字符串实现了所有通用序列操作,可以用for item in <str>或for index, item in enumerate(<str>)语法进行遍历。单个码点(长度为 1 的字符串)可以从左边用0-based index编号引用,也可以从右边用-1-based index编号引用。

字符串可以用<str> + <other str>或<str>.join(<iterable>)拼接,也可以用<str>.split(<separator>)拆分。此外还提供了多种额外的格式化、组装和模板化选项。

str字面量可以用单引号'或双引号"声明。需要时可以使用转义字符\。


>>> single_quoted = 'These allow "double quoting" without "escape" characters.'

>>> double_quoted = "These allow embedded 'single quoting', so you don't have to use an 'escape' character."

多行字符串用'''或"""声明。

>>> triple_quoted = '''Three single quotes or "double quotes" in a row allow for multi-line string literals.
  Line break characters, tabs and other whitespace is fully supported. Remember - The escape "\" character is also available if needed (as can be seen below). 
  
  You\'ll most often encounter multi-line strings as "doc strings" or "doc tests" written just below the first line of a function or class definition.
    They\'re often used with auto documentation ✍ tools.
    '''

可以用 str(<object>)构造函数从其他对象创建字符串或强制转换成字符串:

>>> my_number = 42
>>> str(my_number)
...
"42"

虽然str(<object>)构造函数可以把其他对象强制转换或转换成字符串,但它_不会遍历_或解包对象。这与其他数据类型(如list()、set()、dict()或tuple())的构造函数的行为不同,可能会带来意想不到的结果。

>>> numbers = [1,3,5,7]
>>> str(numbers)
...
'[1,3,5,7]'

str中的码点可以从左边用0-based index编号引用:

creative = '창의적인'

>>> creative[0]
'창'

>>> creative[2]
'적'

>>> creative[3]
'인'

也可以从右边用-1-based index编号引用:

creative = '창의적인'

>>> creative[-4]
'창'

>>> creative[-2]
'적'

>>> creative[-1]
'인'

Python 中没有单独的“字符”或“rune”类型,所以对字符串取下标会生成一个新的str,其长度为 1:


>>> website = "exercism"
>>> type(website[0])
<class 'str'>

>>> len(website[0])
1

>>> website[0] == website[0:1] == 'e'
True

子字符串可以通过_切片语法_选取,使用 <str>[<start>:<stop>:<step>] 生成一个新字符串。结果不包含stop下标。如果没有给出start,起始下标就是 0。如果没有给出stop,stop下标就是字符串的末尾。

moon_and_stars = '🌟🌟🌙🌟🌟⭐'

>>> moon_and_stars[1:4]
'🌟🌙🌟'

>>> moon_and_stars[:3]
'🌟🌟🌙'

>>> moon_and_stars[3:]
'🌟🌟⭐'

>>> moon_and_stars[:-1]
'🌟🌟🌙🌟🌟'

>>> moon_and_stars[:-3]
'🌟🌟🌙'

字符串也可以通过 <str>.split(<separator>) 拆分成更小的字符串,它会返回一个由子字符串组成的list。不带任何参数使用<str>.split()会按空白字符拆分字符串。

>>> cat_ipsum = "Destroy house in 5 seconds command the hooman."
>>> cat_ipsum.split()
...
['Destroy', 'house', 'in', '5', 'seconds', 'command', 'the', 'hooman.']


>>> cat_words = "feline, four-footed, ferocious, furry"
>>> cat_words.split(',')
...
['feline', ' four-footed', ' ferocious', ' furry']


>>> colors = """red,
orange,
green,
purple,
yellow"""

>>> colors.split(',\n')
['red', 'orange', 'green', 'purple', 'yellow']

字符串可以用+运算符拼接。这种方法要少用,因为它的性能不太好,也不容易维护。

language = "Ukrainian"
number = "nine"
word = "дев'ять"

sentence = word + " " + "means" + " " + number + " in " + language + "."

>>> print(sentence)
...
"дев'ять means nine in Ukrainian."

如果需要把list、tuple、set或其他由单个字符串组成的集合合并成一个str,<str>.join(<iterable>)是更好的选择:

# str.join() makes a new string from the iterables elements.
>>> chickens = ["hen", "egg", "rooster"] # Lists are iterable.
>>> ' '.join(chickens)
'hen egg rooster'

# Any string can be used as the joining element.
>>> ' :: '.join(chickens)
'hen :: egg :: rooster'

>>> ' 🌿 '.join(chickens)
'hen 🌿 egg 🌿 rooster'


# Any iterable can be used as input.
>>> flowers = ("rose", "daisy", "carnation")  # Tuples are iterable.
>>> '*-*'.join(flowers)
'rose*-*daisy*-*carnation'

>>> flowers = {"rose", "daisy", "carnation"}  # Sets are iterable, but output order is not guaranteed.
>>> '*-*'.join(flowers)
'rose*-*carnation*-*daisy'

>>> phrase = "This is my string"  # Strings are iterable, but be careful!
>>> '..'.join(phrase)
'T..h..i..s.. ..i..s.. ..m..y.. ..s..t..r..i..n..g'


# Separators are inserted **between** elements, but can be any string (including spaces).
# This can be exploited for interesting effects.
>>> under_words = ['under', 'current', 'sea', 'pin', 'dog', 'lay']
>>> separator = ' ⤴️ under' # Note the leading space, but no trailing space.
>>> separator.join(under_words)
'under ⤴️ undercurrent ⤴️ undersea ⤴️ underpin ⤴️ underdog ⤴️ underlay'

# The separator can be composed different ways, as long as the result is a string.
>>> upper_words = ['upper', 'crust', 'case', 'classmen', 'most', 'cut']
>>> separator = ' 🌟 ' + upper_words[0] # This becomes one string, similar to ' ⤴️ under'.
>>> separator.join(upper_words)
 'upper 🌟 uppercrust 🌟 uppercase 🌟 upperclassmen 🌟 uppermost 🌟 uppercut'

字符串支持所有通用序列操作。单个码点可以在循环中通过for item in <str>遍历。下标可以_连同_元素一起在循环中通过for index, item in enumerate(<str>)遍历。


>>> exercise = 'လေ့ကျင့်'

# Note that there are more code points than perceived glyphs or characters.
# Care should be used when iterating over languages that use
# combining characters, or when dealing with emoji.
>>> for code_point in exercise:
...    print(code_point)
...
လ
ေ
့
က
ျ
င
်
့

# Using enumerate will give both the value and index position of each element.
>>> for index, code_point in enumerate(exercise):
...    print(index, ": ", code_point)
...
0 :  လ
1 :  ေ
2 :  ့
3 :  က
4 :  ျ
5 :  င
6 :  ်
7 :  ့

字符串方法

Python 提供了丰富的字符串方法,可用于搜索、清理、拆分、转换、翻译以及许多其他操作。其中一部分方法会在另一个练习中介绍。

格式化

Python 还提供了丰富的工具来格式化和模板化字符串,并且可以通过 re(正则表达式)、difflib(序列比较) 和 textwrap 模块进行更复杂的文本处理。想找一篇介绍 Python 字符串格式化的好文章,请看 Real Python 上的这篇文章。想了解字符串方法的入门介绍,请看同一网站上的 Strings and Character Data in Python。

相关类型和编码

除了str(一种_文本_序列)之外,Python 还有相应的二进制序列类型,总结在二进制数据服务下,包括bytes(一种_二进制_序列)、bytearray和memoryview,用于高效地存储和处理二进制数据。此外,流允许在不使用回调的情况下通过网络连接收发二进制数据。

通过 GitHub 编辑 该链接会在新窗口或标签页中打开

学习 字符串

练习已锁定

再解锁 4 个练习即可练习 字符串