Go 的 regexp 包提供了对正则表达式的支持。
所接受的正则表达式语法与 Perl、Python 等语言所用的通用语法相同。
搜索模式和输入文本都按 UTF-8 解释。
使用反引号(`)创建字符串时,反斜杠(\) 没有任何特殊含义,也不表示制表符\t或换行符\n这类特殊字符的开头:
"\t\n" // regular string literal with 2 characters: a tab and a newline
`\t\n`// raw string literal with 4 characters: two backslashes, a 't', and an 'n'
因此,写正则表达式时用反引号更合适,因为这样就不需要对反斜杠进行转义:
"\\" // string with a single backslash
`\\` // string with 2 backslashes
RegExp 类型要使用正则表达式,必须先编译字符串模式。
这里的编译是指,把正则表达式的字符串模式转换成一种更易于处理的内部表示。
每个模式只需编译一次,之后就可以多次使用编译后的正则表达式。
regexp.Regexp 类型表示一个已编译的正则表达式。
我们可以使用函数regexp.Compile把字符串模式编译成regexp.Regexp。
如果编译失败,该函数会返回nil和一个错误:
re, err := regexp.Compile(`(a|b)+`)
fmt.Println(re, err) // => (a|b)+ <nil>
re, err = regexp.Compile(`a|b)+`)
fmt.Println(re, err) // => <nil> error parsing regexp: unexpected ): `a|b)+`
函数MustCompile是Compile的便捷替代方案:
re = regexp.MustCompile(`[a-z]+\d*`)
使用这个函数,就不需要处理错误。
MustCompile only should be used when we know for sure the pattern does compile, as otherwise the program will panic.
Regexp有 16 个方法可以匹配正则表达式并定位匹配到的文本。
这些方法的名字都能被下面这个正则表达式匹配:
Find(All)?(String)?(Submatch)?(Index)?
All,该方法会依次匹配整个表达式所有互不重叠的匹配项。String,实参是字符串;否则是字节切片;返回值也会相应地调整。Submatch,返回值是一个切片,标识表达式中依次匹配到的各个子匹配。Index,匹配项和子匹配项由输入字符串中的字节下标对来标识。还有一些方法用于:
总的来说,regexp包定义了 40 多个函数和方法。
下面我们会演示几个方法的用法。
关于这些函数以及其他函数的详细信息,请参阅 API 文档。
MatchString 示例方法MatchString用于报告字符串中是否包含正则表达式的匹配项。
re = regexp.MustCompile(`[a-z]+\d*`)
b = re.MatchString("[a12]") // => true
b = re.MatchString("12abc34(ef)") // => true
b = re.MatchString(" abc!") // => true
b = re.MatchString("123 456") // => false
FindString 示例方法FindString返回一个字符串,其中保存着正则表达式最左侧匹配项的文字。
re = regexp.MustCompile(`[a-z]+\d*`)
s = re.FindString("[a12]") // => "a12"
s = re.FindString("12abc34(ef)") // => "abc34"
s = re.FindString(" abc!") // => "abc"
s = re.FindString("123 456") // => ""
FindStringSubmatch 示例方法FindStringSubmatch返回一个字符串切片,其中保存着正则表达式最左侧匹配项的文字,以及它的子表达式(如果有)匹配到的文字。
这可以用来找出与捕获组匹配的字符串。
返回值为nil表示没有匹配。
re = regexp.MustCompile(`[a-z]+(\d*)`)
sl = re.FindStringSubmatch("[a12]") // => []string{"a12","12"}
sl = re.FindStringSubmatch("12abc34(ef)") // => []string{"abc34","34"}
sl = re.FindStringSubmatch(" abc!") // => []string{"abc",""}
sl = re.FindStringSubmatch("123 456") // => <nil>
ReplaceAllString 示例方法re.ReplaceAllString(src,repl)返回src的副本,并把正则表达式re的匹配项替换为替换字符串repl。
re = regexp.MustCompile(`[a-z]+\d*`)
s = re.ReplaceAllString("[a12]", "X") // => "[X]"
s = re.ReplaceAllString("12abc34(ef)", "X") // => "12X(X)"
s = re.ReplaceAllString(" abc!", "X") // => " X!"
s = re.ReplaceAllString("123 456", "X") // => "123 456"
Split 示例方法re.Split(s,n)会把文本s按该表达式切分成子字符串,并返回这些表达式匹配项之间的子字符串所组成的切片。
数量n决定最多返回多少个子字符串。
如果n<0,该方法返回所有子字符串。
re = regexp.MustCompile(`[a-z]+\d*`)
sl = re.Split("[a12]", -1) // => []string{"[","]"}
sl = re.Split("12abc34(ef)", 2) // => []string{"12","(ef)"}
sl = re.Split(" abc!", -1) // => []string{" ","!"}
sl = re.Split("123 456", -1) // => []string{"123 456"}
这道练习的主题是解析日志文件。
在最近的一次安全审查之后,你被要求清理组织归档的日志文件。
传给这些函数的所有字符串都保证非空,并且首尾没有空格。
你需要大致了解归档里有多少日志行不符合当前标准。 你认为只要一个简单的测试就能判断某行日志是否有效。 一行日志要被认定为有效,必须以以下字符串之一开头:
实现IsValidLine函数:如果字符串无效就返回false,否则返回true。
IsValidLine("[ERR] A good error here")
// => true
IsValidLine("Any old [ERR] text")
// => false
IsValidLine("[BOB] Any old text")
// => false
一个新团队加入了组织,你发现他们的日志文件用一种奇怪的分隔符来分隔“字段”。 他们不用像冒号 ":" 这样合理的东西,而是用 "<--->" 或 "<=>" 这样的字符串(因为这样更好看)。实际上,只要第一个字符是 "<"、最后一个字符是 ">",中间是 "~"、"*"、"=" 和 "-" 的任意组合,这样的字符串都可以。
实现SplitLogLine函数,它接收一行日志,返回一个字符串数组,其中每个字符串都包含一个字段。
SplitLogLine("section 1<*>section 2<~~~>section 3")
// => []string{"section 1", "section 2", "section 3"},
password的行数团队需要了解引号文本中对密码的提及,以便人工检查。
实现CountQuotedPasswords函数,用来大致估计这项人工检查可能的规模。
找出这样的日志行:字符串 "password" 被引号包围,其中字母的大小写可以是任意组合。 你要考虑到引号内 "password" 前后可能还有其他内容。 每行最多包含两个引号。
传给这个函数的行可能符合第 1 个任务定义的有效性,也可能不符合。 无论是否有效,我们都以同样的方式处理它们。
lines := []string{
`[INF] passWord`, // contains 'password' but not surrounded by quotation marks
`"passWord"`, // count this one
`[INF] User saw error message "Unexpected Error" on page load.`, // does not contain 'password'
`[INF] The message "Please reset your password" was ignored by the user`, // count this one
}
// => 2
你发现日志的某些上游处理会在整个日志中到处散布 "end-of-line" 加上行号的文本(中间没有空格)。
实现RemoveEndOfLineText函数,它接收一个字符串,移除其中的 end-of-line 文本,返回一个“干净”的字符串。
不包含 end-of-line 文本的行应原样返回。
只需移除 end-of-line 字符串。 不要试图调整空白字符。
RemoveEndOfLineText("[INF] end-of-line23033 Network Failure end-of-line27")
// => "[INF] Network Failure "
你注意到有些日志行包含提到用户的句子。
这些句子总是包含字符串"User",后面跟着一个或多个空格字符,然后是一个用户名。
你决定给这样的行打上标签。
实现一个函数TagWithUserName来处理日志行:
"User "的行保持不变。"User "的行,在行首加上[USR],后面紧跟用户名。例如:
result := TagWithUserName([]string{
"[WRN] User James123 has exceeded storage space.",
"[WRN] Host down. User Michelle4 lost connection.",
"[INF] Users can login again after 23:00.",
"[DBG] We need to check that user names are at least 6 chars long.",
})
// => []string {
// "[USR] James123 [WRN] User James123 has exceeded storage space.",
// "[USR] Michelle4 [WRN] Host down. User Michelle4 lost connection.",
// "[INF] Users can login again after 23:00.",
// "[DBG] We need to check that user names are at least 6 chars long."
// }
你可以假设: