Python 中的str是不可變序列,由 Unicode 碼位組成。這些碼位可能包含字母、變音符號、定位字元、數字、貨幣符號、表情符號、標點符號、空格與換行字元等等。
若想深入了解字串究竟編碼了哪些資訊(或者說,「電腦怎麼知道要如何把 0 與 1 轉換成字母?」),這篇部落格文章至今仍然很有幫助。Python 官方文件也提供了非常詳細的 unicode HOWTO,討論 Python 在str、bytes與re模組中對 Unicode 規格的支援、地區設定的各種考量,以及一些編碼與轉譯的常見問題。
字串支援所有常見序列操作,也可以使用 for item in <str> 或 for index, item in enumerate(<str>) 語法來疊代。個別的碼位(長度為 1 的字串)可以從左邊以 0-based index 的編號取用,或從右邊以 -1-based index 的編號取用。
字串可以用 <str> + <other str> 或 <str>.join(<iterable>) 串接,並透過 <str>.split(<separator>) 分割。字串還提供多種額外的格式化、組裝與樣板選項。
str 字面值可以使用單引號'或雙引號"來宣告。需要時也可以使用跳脫字元\。
>>> single_quoted = 'These allow "double quoting" without "escape" characters.'
>>> double_quoted = "These allow embedded 'single quoting', so you don't have to use an 'escape' character."
多行字串則使用'''或"""來宣告。
>>> triple_quoted = '''Three single quotes or "double quotes" in a row allow for multi-line string literals.
Line break characters, tabs and other whitespace is fully supported. Remember - The escape "\" character is also available if needed (as can be seen below).
You\'ll most often encounter multi-line strings as "doc strings" or "doc tests" written just below the first line of a function or class definition.
They\'re often used with auto documentation ✍ tools.
'''
可以用 str(<object>) 建構子從其他物件建立或強制轉換出字串:
>>> my_number = 42
>>> str(my_number)
...
"42"
雖然 str(<object>) 建構子可以用來強制轉換或轉換成字串,但它 _不會疊代_或拆解物件。這和 list()、set()、dict() 或 tuple() 等其他資料型態的建構子行為不同,可能會帶來令人意外的結果。
>>> numbers = [1,3,5,7]
>>> str(numbers)
...
'[1,3,5,7]'
str中的碼位可以從左邊以 0-based index 的編號取用:
creative = '창의적인'
>>> creative[0]
'창'
>>> creative[2]
'적'
>>> creative[3]
'인'
也可以從右邊索引,從 -1-based index 開始:
creative = '창의적인'
>>> creative[-4]
'창'
>>> creative[-2]
'적'
>>> creative[-1]
'인'
Python 中沒有獨立的「character」或「rune」型態,所以對字串做索引會產生一個長度為 1的新str:
>>> website = "exercism"
>>> type(website[0])
<class 'str'>
>>> len(website[0])
1
>>> website[0] == website[0:1] == 'e'
True
子字串可以透過 切片標記法 選取,使用 <str>[<start>:<stop>:<step>] 產生新的字串。結果不包含 stop 索引。如果沒有指定 start,起始索引會是 0。如果沒有指定 stop,stop 索引就會是字串的結尾。
moon_and_stars = '🌟🌟🌙🌟🌟⭐'
>>> moon_and_stars[1:4]
'🌟🌙🌟'
>>> moon_and_stars[:3]
'🌟🌟🌙'
>>> moon_and_stars[3:]
'🌟🌟⭐'
>>> moon_and_stars[:-1]
'🌟🌟🌙🌟🌟'
>>> moon_and_stars[:-3]
'🌟🌟🌙'
字串也可以透過 <str>.split(<separator>) 拆成更小的字串,它會回傳一份由子字串組成的 list。使用不帶任何引數的 <str>.split() 會以空白字元分割字串。
>>> cat_ipsum = "Destroy house in 5 seconds command the hooman."
>>> cat_ipsum.split()
...
['Destroy', 'house', 'in', '5', 'seconds', 'command', 'the', 'hooman.']
>>> cat_words = "feline, four-footed, ferocious, furry"
>>> cat_words.split(',')
...
['feline', ' four-footed', ' ferocious', ' furry']
>>> colors = """red,
orange,
green,
purple,
yellow"""
>>> colors.split(',\n')
['red', 'orange', 'green', 'purple', 'yellow']
字串可以使用 + 運算子串接。這種方法應該少用,因為它的效能不太好,也不容易維護。
language = "Ukrainian"
number = "nine"
word = "дев'ять"
sentence = word + " " + "means" + " " + number + " in " + language + "."
>>> print(sentence)
...
"дев'ять means nine in Ukrainian."
如果需要把 list、tuple、set 或其他由個別字串組成的集合合併成單一的 str,<str>.join(<iterable>) 是更好的選擇:
# str.join() makes a new string from the iterables elements.
>>> chickens = ["hen", "egg", "rooster"] # Lists are iterable.
>>> ' '.join(chickens)
'hen egg rooster'
# Any string can be used as the joining element.
>>> ' :: '.join(chickens)
'hen :: egg :: rooster'
>>> ' 🌿 '.join(chickens)
'hen 🌿 egg 🌿 rooster'
# Any iterable can be used as input.
>>> flowers = ("rose", "daisy", "carnation") # Tuples are iterable.
>>> '*-*'.join(flowers)
'rose*-*daisy*-*carnation'
>>> flowers = {"rose", "daisy", "carnation"} # Sets are iterable, but output order is not guaranteed.
>>> '*-*'.join(flowers)
'rose*-*carnation*-*daisy'
>>> phrase = "This is my string" # Strings are iterable, but be careful!
>>> '..'.join(phrase)
'T..h..i..s.. ..i..s.. ..m..y.. ..s..t..r..i..n..g'
# Separators are inserted **between** elements, but can be any string (including spaces).
# This can be exploited for interesting effects.
>>> under_words = ['under', 'current', 'sea', 'pin', 'dog', 'lay']
>>> separator = ' ⤴️ under' # Note the leading space, but no trailing space.
>>> separator.join(under_words)
'under ⤴️ undercurrent ⤴️ undersea ⤴️ underpin ⤴️ underdog ⤴️ underlay'
# The separator can be composed different ways, as long as the result is a string.
>>> upper_words = ['upper', 'crust', 'case', 'classmen', 'most', 'cut']
>>> separator = ' 🌟 ' + upper_words[0] # This becomes one string, similar to ' ⤴️ under'.
>>> separator.join(upper_words)
'upper 🌟 uppercrust 🌟 uppercase 🌟 upperclassmen 🌟 uppermost 🌟 uppercut'
字串支援所有常見序列操作。個別的碼位可以在迴圈中用 for item in <str> 疊代。索引_與_項目可以在迴圈中用 for index, item in enumerate(<str>) 一起疊代。
>>> exercise = 'လေ့ကျင့်'
# Note that there are more code points than perceived glyphs or characters.
# Care should be used when iterating over languages that use
# combining characters, or when dealing with emoji.
>>> for code_point in exercise:
... print(code_point)
...
လ
ေ
့
က
ျ
င
်
့
# Using enumerate will give both the value and index position of each element.
>>> for index, code_point in enumerate(exercise):
... print(index, ": ", code_point)
...
0 : လ
1 : ေ
2 : ့
3 : က
4 : ျ
5 : င
6 : ်
7 : ့
Python 提供了豐富的字串方法,可以協助搜尋、清理、分割、轉換、轉譯,以及許多其他操作。其中一部分方法會在另一個練習中介紹。
Python 也提供了豐富的工具,可用來格式化與套用樣板字串,也可以透過 re(正規表達式)、difflib(序列比對) 和 textwrap 模組進行更精密的文字處理。想好好認識 Python 的字串格式化,請參考 Real Python 的這篇文章。想認識字串方法,請參考同一網站的 Strings and Character Data in Python。
除了 str(一種_文字_序列),Python 還有對應的二進位序列型態,彙整在二進位資料服務之下:bytes(一種_二進位_序列)、bytearray 和 memoryview,可用來有效率地儲存與處理二進位資料。此外,Streams 讓你能透過網路連線收發二進位資料,而不需要使用回呼函式。