Go 中的rune型別是int32的別名。
由於底層型別是int32,rune型別所存放的是帶正負號的 32 位元整數值。
不過,rune型別和int32型別不同:rune型別裡存放的整數值代表單一的 Unicode 字元。
Unicode 是 ASCII 的超集,它為每個字元指定一個獨一無二的編號來表示該字元。 這個獨一無二的編號就稱為 Unicode 碼位。 Unicode 的目標是把世界上所有的字元,包括各種字母、數字、符號,甚至 emoji,都以 Unicode 碼位來表示。
在 Go 中,rune型別代表單一的 Unicode 碼位。
下表列出一些 Unicode 字元範例,以及它們的 Unicode 碼位和十進位值:
| Unicode 字元 | Unicode 碼位 | 十進位值 |
|---|---|---|
| 0 | U+0030 |
48 |
| A | U+0041 |
65 |
| a | U+0061 |
97 |
| ¿ | U+00BF |
191 |
| π | U+03C0 |
960 |
| 🧠 | U+1F9E0 |
129504 |
UTF-8 是一種可變寬度的字元編碼,用來把每個 Unicode 碼位編碼成 1、2、3 或 4 個位元組。
由於一個 Unicode 碼位最多可以編碼成 4 個位元組,rune型別必須能容納最多 4 個位元組的資料。
這就是為什麼rune型別是int32的別名,因為int32型別正好能容納最多 4 個位元組的資料。
Go 的原始碼檔案是以 UTF-8 編碼。
rune型別的變數是把字元放在單引號裡來宣告:
myRune := '¿'
由於rune只是int32的別名,印出 rune 的型別會得到int32:
myRune := '¿'
fmt.Printf("myRune type: %T\n", myRune)
// Output: myRune type: int32
同樣地,印出 rune 的值會得到它的整數(十進位)值:
myRune := '¿'
fmt.Printf("myRune value: %v\n", myRune)
// Output: myRune value: 191
要印出 rune 所代表的 Unicode 字元,請使用%c格式化動詞:
myRune := '¿'
fmt.Printf("myRune Unicode character: %c\n", myRune)
// Output: myRune Unicode character: ¿
要印出 rune 所代表的 Unicode 碼位,請使用%U格式化動詞:
myRune := '¿'
fmt.Printf("myRune Unicode code point: %U\n", myRune)
// Output: myRune Unicode code point: U+00BF
除了用單引號把字元包起來定義 rune 之外,你也可以直接指定十六進位或十進位的數字:
myRune := rune(0xbf)
myRune = 191
fmt.Printf("myRune Unicode character: %c\n", myRune)
// Output: myRune Unicode character: ¿
Go 的字串是以 UTF-8 編碼,這表示字串裡包含的是 Unicode 字元。 字串中的字元依它所代表的 Unicode 字元,儲存和編碼成 1、2、3 或 4 個位元組。
在 Go 中,序列是用切片來表示,而這些切片可以用 range 來疊代。
當我們疊代一個字串時,Go 會把字串轉換成一連串的 rune,每個 rune 佔 4 個位元組(還記得嗎?rune 型別是int32的別名!)
雖然字串只是位元組的切片,但range關鍵字疊代的是字串的 rune,而不是它的位元組。
在這個範例中,index變數代表目前 rune 位元組序列的起始索引,char變數則代表目前的 rune:
myString := "❗hello"
for index, char := range myString {
fmt.Printf("Index: %d\tCharacter: %c\t\tCode Point: %U\n", index, char, char)
}
// Output:
// Index: 0 Character: ❗ Code Point: U+2757
// Index: 3 Character: h Code Point: U+0068
// Index: 4 Character: e Code Point: U+0065
// Index: 5 Character: l Code Point: U+006C
// Index: 6 Character: l Code Point: U+006C
// Index: 7 Character: o Code Point: U+006F
由於 rune 可以儲存成 1、2、3 或 4 個位元組,字串的長度不一定等於字串中的字元數。
用內建的len函式可以取得字串以位元組計的長度,用utf8.RuneCountInString函式則可以取得字串中的 rune 數量:
import "unicode/utf8"
myString := "❗hello"
stringLength := len(myString)
numberOfRunes := utf8.RuneCountInString(myString)
fmt.Printf("myString - Length: %d - Runes: %d\n", stringLength, numberOfRunes)
// Output: myString - Length: 8 - Runes: 6
rune 切片可以型別轉換成字串:
myRuneSlice := []rune{'e', 'x', 'e', 'r', 'c', 'i', 's', 'm'}
myString := string(myRuneSlice)
fmt.Println(myString)
// Output: exercism
同樣地,字串也可以型別轉換成 rune 切片。 記得,沒有格式化動詞時,印出 rune 會得到它的整數(十進位)值:
myString := "exercism"
myRuneSlice := []rune(myString)
fmt.Println(myRuneSlice)
// Output: [101 120 101 114 99 105 115 109]