Go 里的rune类型是int32的别名。由于底层是int32类型,rune类型保存的是一个有符号的 32 位整数值。不过,与int32类型不同的是,rune类型中存储的整数值表示单个 Unicode 字符。
Unicode 是 ASCII 的超集,它给每个字符分配一个唯一的编号来表示字符。这个唯一的编号称为 Unicode 码点。Unicode 的目标是把世界上所有的字符都表示为 Unicode 码点,包括各种字母表、数字、符号,甚至 emoji。
在 Go 中,rune类型表示单个 Unicode 码点。
下表列出了一些示例 Unicode 字符及其 Unicode 码点和十进制值:
| Unicode 字符 | Unicode 码点 | 十进制值 |
|---|---|---|
| 0 | U+0030 |
48 |
| A | U+0041 |
65 |
| a | U+0061 |
97 |
| ¿ | U+00BF |
191 |
| π | U+03C0 |
960 |
| 🧠 | U+1F9E0 |
129504 |
UTF-8 是一种变长字符编码,用于把每个 Unicode 码点编码为 1、2、3 或 4 个字节。由于一个 Unicode 码点最多可以编码为 4 个字节,rune类型需要能够容纳最多 4 个字节的数据。这就是为什么rune类型是int32的别名,因为int32类型能够容纳最多 4 个字节的数据。
Go 源代码文件使用 UTF-8 编码。
rune类型的变量通过把字符放在单引号里来声明:
myRune := '¿'
由于rune只是int32的别名,打印 rune 的类型会得到int32:
myRune := '¿'
fmt.Printf("myRune type: %T\n", myRune)
// Output: myRune type: int32
同样,打印 rune 的值会得到它的整数(十进制)值:
myRune := '¿'
fmt.Printf("myRune value: %v\n", myRune)
// Output: myRune value: 191
要打印 rune 所表示的 Unicode 字符,使用%c格式化动词:
myRune := '¿'
fmt.Printf("myRune Unicode character: %c\n", myRune)
// Output: myRune Unicode character: ¿
要打印 rune 所表示的 Unicode 码点,使用%U格式化动词:
myRune := '¿'
fmt.Printf("myRune Unicode code point: %U\n", myRune)
// Output: myRune Unicode code point: U+00BF
除了用单引号把字符括起来定义 rune,你还可以直接指定十六进制或十进制数字:
myRune := rune(0xbf)
myRune = 191
fmt.Printf("myRune Unicode character: %c\n", myRune)
// Output: myRune Unicode character: ¿
Go 中的字符串使用 UTF-8 编码,这意味着它们包含 Unicode 字符。字符串中的字符根据其所表示的 Unicode 字符,以 1、2、3 或 4 个字节存储和编码。
在 Go 中,切片用于表示序列,并且可以用 range 迭代这些切片。当我们迭代一个字符串时,Go 会把字符串转换成一系列 rune,每个 rune 占 4 个字节(记住,rune 类型是int32的别名!)
尽管字符串只是一个字节切片,但range关键字迭代的是字符串的 rune,而不是它的字节。
在这个例子中,index变量表示当前 rune 字节序列的起始下标,char变量表示当前的 rune:
myString := "❗hello"
for index, char := range myString {
fmt.Printf("Index: %d\tCharacter: %c\t\tCode Point: %U\n", index, char, char)
}
// Output:
// Index: 0 Character: ❗ Code Point: U+2757
// Index: 3 Character: h Code Point: U+0068
// Index: 4 Character: e Code Point: U+0065
// Index: 5 Character: l Code Point: U+006C
// Index: 6 Character: l Code Point: U+006C
// Index: 7 Character: o Code Point: U+006F
由于 rune 可以以 1、2、3 或 4 个字节存储,字符串的长度不一定等于字符串中的字符数。使用内置的len函数获取字符串的字节长度,使用utf8.RuneCountInString函数获取字符串中 rune 的数量:
import "unicode/utf8"
myString := "❗hello"
stringLength := len(myString)
numberOfRunes := utf8.RuneCountInString(myString)
fmt.Printf("myString - Length: %d - Runes: %d\n", stringLength, numberOfRunes)
// Output: myString - Length: 8 - Runes: 6
rune 切片可以类型转换为字符串:
myRuneSlice := []rune{'e', 'x', 'e', 'r', 'c', 'i', 's', 'm'}
myString := string(myRuneSlice)
fmt.Println(myString)
// Output: exercism
同样,字符串也可以类型转换为 rune 切片。记住,不使用格式化动词时,打印 rune 会得到它的整数(十进制)值:
myString := "exercism"
myRuneSlice := []rune(myString)
fmt.Println(myRuneSlice)
// Output: [101 120 101 114 99 105 115 109]