字串

字串 在 Python

83 個練習

關於 字串

Python 中的str是不可變序列,由 Unicode 碼位組成。這些碼位可能包含字母、變音符號、定位字元、數字、貨幣符號、表情符號、標點符號、空格與換行字元等等。

若想深入了解字串究竟編碼了哪些資訊(或者說,「電腦怎麼知道要如何把 0 與 1 轉換成字母?」),這篇部落格文章至今仍然很有幫助。Python 官方文件也提供了非常詳細的 unicode HOWTO,討論 Python 在str、bytes與re模組中對 Unicode 規格的支援、地區設定的各種考量,以及一些編碼與轉譯的常見問題。

字串支援所有常見序列操作,也可以使用 for item in <str> 或 for index, item in enumerate(<str>) 語法來疊代。個別的碼位(長度為 1 的字串)可以從左邊以 0-based index 的編號取用,或從右邊以 -1-based index 的編號取用。

字串可以用 <str> + <other str> 或 <str>.join(<iterable>) 串接,並透過 <str>.split(<separator>) 分割。字串還提供多種額外的格式化、組裝與樣板選項。

str 字面值可以使用單引號'或雙引號"來宣告。需要時也可以使用跳脫字元\。


>>> single_quoted = 'These allow "double quoting" without "escape" characters.'

>>> double_quoted = "These allow embedded 'single quoting', so you don't have to use an 'escape' character."

多行字串則使用'''或"""來宣告。

>>> triple_quoted = '''Three single quotes or "double quotes" in a row allow for multi-line string literals.
  Line break characters, tabs and other whitespace is fully supported. Remember - The escape "\" character is also available if needed (as can be seen below). 
  
  You\'ll most often encounter multi-line strings as "doc strings" or "doc tests" written just below the first line of a function or class definition.
    They\'re often used with auto documentation ✍ tools.
    '''

可以用 str(<object>) 建構子從其他物件建立或強制轉換出字串:

>>> my_number = 42
>>> str(my_number)
...
"42"

雖然 str(<object>) 建構子可以用來強制轉換或轉換成字串,但它 _不會疊代_或拆解物件。這和 list()、set()、dict() 或 tuple() 等其他資料型態的建構子行為不同,可能會帶來令人意外的結果。

>>> numbers = [1,3,5,7]
>>> str(numbers)
...
'[1,3,5,7]'

str中的碼位可以從左邊以 0-based index 的編號取用:

creative = '창의적인'

>>> creative[0]
'창'

>>> creative[2]
'적'

>>> creative[3]
'인'

也可以從右邊索引,從 -1-based index 開始:

creative = '창의적인'

>>> creative[-4]
'창'

>>> creative[-2]
'적'

>>> creative[-1]
'인'

Python 中沒有獨立的「character」或「rune」型態,所以對字串做索引會產生一個長度為 1的新str:


>>> website = "exercism"
>>> type(website[0])
<class 'str'>

>>> len(website[0])
1

>>> website[0] == website[0:1] == 'e'
True

子字串可以透過 切片標記法 選取,使用 <str>[<start>:<stop>:<step>] 產生新的字串。結果不包含 stop 索引。如果沒有指定 start,起始索引會是 0。如果沒有指定 stop,stop 索引就會是字串的結尾。

moon_and_stars = '🌟🌟🌙🌟🌟⭐'

>>> moon_and_stars[1:4]
'🌟🌙🌟'

>>> moon_and_stars[:3]
'🌟🌟🌙'

>>> moon_and_stars[3:]
'🌟🌟⭐'

>>> moon_and_stars[:-1]
'🌟🌟🌙🌟🌟'

>>> moon_and_stars[:-3]
'🌟🌟🌙'

字串也可以透過 <str>.split(<separator>) 拆成更小的字串,它會回傳一份由子字串組成的 list。使用不帶任何引數的 <str>.split() 會以空白字元分割字串。

>>> cat_ipsum = "Destroy house in 5 seconds command the hooman."
>>> cat_ipsum.split()
...
['Destroy', 'house', 'in', '5', 'seconds', 'command', 'the', 'hooman.']


>>> cat_words = "feline, four-footed, ferocious, furry"
>>> cat_words.split(',')
...
['feline', ' four-footed', ' ferocious', ' furry']


>>> colors = """red,
orange,
green,
purple,
yellow"""

>>> colors.split(',\n')
['red', 'orange', 'green', 'purple', 'yellow']

字串可以使用 + 運算子串接。這種方法應該少用,因為它的效能不太好,也不容易維護。

language = "Ukrainian"
number = "nine"
word = "дев'ять"

sentence = word + " " + "means" + " " + number + " in " + language + "."

>>> print(sentence)
...
"дев'ять means nine in Ukrainian."

如果需要把 list、tuple、set 或其他由個別字串組成的集合合併成單一的 str,<str>.join(<iterable>) 是更好的選擇:

# str.join() makes a new string from the iterables elements.
>>> chickens = ["hen", "egg", "rooster"] # Lists are iterable.
>>> ' '.join(chickens)
'hen egg rooster'

# Any string can be used as the joining element.
>>> ' :: '.join(chickens)
'hen :: egg :: rooster'

>>> ' 🌿 '.join(chickens)
'hen 🌿 egg 🌿 rooster'


# Any iterable can be used as input.
>>> flowers = ("rose", "daisy", "carnation")  # Tuples are iterable.
>>> '*-*'.join(flowers)
'rose*-*daisy*-*carnation'

>>> flowers = {"rose", "daisy", "carnation"}  # Sets are iterable, but output order is not guaranteed.
>>> '*-*'.join(flowers)
'rose*-*carnation*-*daisy'

>>> phrase = "This is my string"  # Strings are iterable, but be careful!
>>> '..'.join(phrase)
'T..h..i..s.. ..i..s.. ..m..y.. ..s..t..r..i..n..g'


# Separators are inserted **between** elements, but can be any string (including spaces).
# This can be exploited for interesting effects.
>>> under_words = ['under', 'current', 'sea', 'pin', 'dog', 'lay']
>>> separator = ' ⤴️ under' # Note the leading space, but no trailing space.
>>> separator.join(under_words)
'under ⤴️ undercurrent ⤴️ undersea ⤴️ underpin ⤴️ underdog ⤴️ underlay'

# The separator can be composed different ways, as long as the result is a string.
>>> upper_words = ['upper', 'crust', 'case', 'classmen', 'most', 'cut']
>>> separator = ' 🌟 ' + upper_words[0] # This becomes one string, similar to ' ⤴️ under'.
>>> separator.join(upper_words)
 'upper 🌟 uppercrust 🌟 uppercase 🌟 upperclassmen 🌟 uppermost 🌟 uppercut'

字串支援所有常見序列操作。個別的碼位可以在迴圈中用 for item in <str> 疊代。索引_與_項目可以在迴圈中用 for index, item in enumerate(<str>) 一起疊代。


>>> exercise = 'လေ့ကျင့်'

# Note that there are more code points than perceived glyphs or characters.
# Care should be used when iterating over languages that use
# combining characters, or when dealing with emoji.
>>> for code_point in exercise:
...    print(code_point)
...
လ
ေ
့
က
ျ
င
်
့

# Using enumerate will give both the value and index position of each element.
>>> for index, code_point in enumerate(exercise):
...    print(index, ": ", code_point)
...
0 :  လ
1 :  ေ
2 :  ့
3 :  က
4 :  ျ
5 :  င
6 :  ်
7 :  ့

字串方法

Python 提供了豐富的字串方法,可以協助搜尋、清理、分割、轉換、轉譯,以及許多其他操作。其中一部分方法會在另一個練習中介紹。

格式化

Python 也提供了豐富的工具,可用來格式化與套用樣板字串,也可以透過 re(正規表達式)、difflib(序列比對) 和 textwrap 模組進行更精密的文字處理。想好好認識 Python 的字串格式化,請參考 Real Python 的這篇文章。想認識字串方法,請參考同一網站的 Strings and Character Data in Python。

相關型態與編碼

除了 str(一種_文字_序列),Python 還有對應的二進位序列型態,彙整在二進位資料服務之下:bytes(一種_二進位_序列)、bytearray 和 memoryview,可用來有效率地儲存與處理二進位資料。此外,Streams 讓你能透過網路連線收發二進位資料,而不需要使用回呼函式。

透過 GitHub 編輯 連結會在新視窗或分頁中開啟

學習 字串

練習已鎖定

再解鎖 4 個練習,就能練習 字串