Ru

Runes 在 Go

3 個練習

關於 Runes

Go 中的rune型別是int32的別名。 由於底層型別是int32,rune型別所存放的是帶正負號的 32 位元整數值。 不過,rune型別和int32型別不同:rune型別裡存放的整數值代表單一的 Unicode 字元。

Unicode 與 Unicode 碼位

Unicode 是 ASCII 的超集,它為每個字元指定一個獨一無二的編號來表示該字元。 這個獨一無二的編號就稱為 Unicode 碼位。 Unicode 的目標是把世界上所有的字元,包括各種字母、數字、符號,甚至 emoji,都以 Unicode 碼位來表示。

在 Go 中,rune型別代表單一的 Unicode 碼位。

下表列出一些 Unicode 字元範例,以及它們的 Unicode 碼位和十進位值:

Unicode 字元 Unicode 碼位 十進位值
0 U+0030 48
A U+0041 65
a U+0061 97
¿ U+00BF 191
π U+03C0 960
🧠 U+1F9E0 129504

UTF-8

UTF-8 是一種可變寬度的字元編碼,用來把每個 Unicode 碼位編碼成 1、2、3 或 4 個位元組。 由於一個 Unicode 碼位最多可以編碼成 4 個位元組,rune型別必須能容納最多 4 個位元組的資料。 這就是為什麼rune型別是int32的別名,因為int32型別正好能容納最多 4 個位元組的資料。

Go 的原始碼檔案是以 UTF-8 編碼。

使用 rune

rune型別的變數是把字元放在單引號裡來宣告:

myRune := '¿'

由於rune只是int32的別名,印出 rune 的型別會得到int32:

myRune := '¿'
fmt.Printf("myRune type: %T\n", myRune)
// Output: myRune type: int32

同樣地,印出 rune 的值會得到它的整數(十進位)值:

myRune := '¿'
fmt.Printf("myRune value: %v\n", myRune)
// Output: myRune value: 191

要印出 rune 所代表的 Unicode 字元,請使用%c格式化動詞:

myRune := '¿'
fmt.Printf("myRune Unicode character: %c\n", myRune)
// Output: myRune Unicode character: ¿

要印出 rune 所代表的 Unicode 碼位,請使用%U格式化動詞:

myRune := '¿'
fmt.Printf("myRune Unicode code point: %U\n", myRune)
// Output: myRune Unicode code point: U+00BF

除了用單引號把字元包起來定義 rune 之外,你也可以直接指定十六進位或十進位的數字:

myRune := rune(0xbf)
myRune = 191
fmt.Printf("myRune Unicode character: %c\n", myRune)
// Output: myRune Unicode character: ¿

rune 與字串

Go 的字串是以 UTF-8 編碼,這表示字串裡包含的是 Unicode 字元。 字串中的字元依它所代表的 Unicode 字元,儲存和編碼成 1、2、3 或 4 個位元組。

在 Go 中,序列是用切片來表示,而這些切片可以用 range 來疊代。 當我們疊代一個字串時,Go 會把字串轉換成一連串的 rune,每個 rune 佔 4 個位元組(還記得嗎?rune 型別是int32的別名!)

雖然字串只是位元組的切片,但range關鍵字疊代的是字串的 rune,而不是它的位元組。

在這個範例中,index變數代表目前 rune 位元組序列的起始索引,char變數則代表目前的 rune:

myString := "❗hello"
for index, char := range myString {
  fmt.Printf("Index: %d\tCharacter: %c\t\tCode Point: %U\n", index, char, char)
}
// Output:
// Index: 0	Character: ❗		Code Point: U+2757
// Index: 3	Character: h		Code Point: U+0068
// Index: 4	Character: e		Code Point: U+0065
// Index: 5	Character: l		Code Point: U+006C
// Index: 6	Character: l		Code Point: U+006C
// Index: 7	Character: o		Code Point: U+006F

由於 rune 可以儲存成 1、2、3 或 4 個位元組,字串的長度不一定等於字串中的字元數。 用內建的len函式可以取得字串以位元組計的長度,用utf8.RuneCountInString函式則可以取得字串中的 rune 數量:

import "unicode/utf8"

myString := "❗hello"
stringLength := len(myString)
numberOfRunes := utf8.RuneCountInString(myString)

fmt.Printf("myString - Length: %d - Runes: %d\n", stringLength, numberOfRunes)
// Output: myString - Length: 8 - Runes: 6

轉換 rune 的型別

rune 切片可以型別轉換成字串:

myRuneSlice := []rune{'e', 'x', 'e', 'r', 'c', 'i', 's', 'm'}
myString := string(myRuneSlice)
fmt.Println(myString)
// Output: exercism

同樣地,字串也可以型別轉換成 rune 切片。 記得,沒有格式化動詞時,印出 rune 會得到它的整數(十進位)值:

myString := "exercism"
myRuneSlice := []rune(myString)
fmt.Println(myRuneSlice)
// Output: [101 120 101 114 99 105 115 109]
透過 GitHub 編輯 連結會在新視窗或分頁中開啟

學習 Runes

練習已鎖定

再解鎖 2 個練習,就能練習 Runes