正则表达式是一种非常灵活的字符串_模式匹配_方式,它使用为此专门设计的领域特定语言(DSL)。
与另外几种编程语言一样,Julia 并不尝试实现自己的正则表达式库。它改为封装了流行的 PCRE2 库,因此提供的正则表达式语法与(例如)Gleam 完全相同,与 Javascript 则非常接近。
这份 Julia 教学大纲假定你已经熟悉基本的正则表达式语法。 我们只专注于 Julia 特有的功能。
下面列出了一些可以帮你复习正则表达式知识的资源。
Julia 的正则表达式接口在手册中有说明。
Julia 中的正则表达式,就是在开头的"之前加上r的字符串。所有基本功能都属于标准库的一部分。
事实上,Strings概念中已经讨论过的许多函数,默认就是为正则表达式搜索而设计的,例如 occursin()。
julia> re = r"test$"
r"test$"
julia> typeof(re)
Regex
# Does a string end with "test"?
julia> occursin(re, "this is a test")
true
julia> occursin(re, "these are tests")
false
修饰符字符可以跟在结尾引号之后,例如i表示不区分大小写的匹配。
julia> occursin(r"test", "Testing")
false
julia> occursin(r"test"i, "Testing")
true
通常,我们想知道_匹配到的究竟是什么_。做法是在正则表达式里用括号加入捕获组,然后使用match()函数。
julia> m = match(r"(\d+g) .* (\d+ml)", "dissolve 25g sugar in 200ml water")
RegexMatch("25g sugar in 200ml", 1="25g", 2="200ml")
julia> m.captures
2-element Vector{Union{Nothing, SubString{String}}}:
"25g"
"200ml"
# how many matches?
julia> length(m.captures)
2
# what matched?
julia> m[1], m[2]
("25g", "200ml")
# Starting positions of the matches (character index)
julia> m.offsets
2-element Vector{Int64}:
10
23
当然,匹配也可能失败。这时结果会是特殊值Nothing,而不是RegexMatch,所以要准备好判断这种情况。
# failed match
m = match(r"(not here)", "dissolve 25g sugar in 200ml water")
julia> typeof(m)
Nothing
julia> isnothing(m)
true
虽然match默认从字符串开头开始匹配,我们也可以指定一个偏移量n,忽略前n个字符。
# capture first match
julia> m = match(r"(\wat)", "cat, sat, mat")
RegexMatch("cat", 1="cat")
# ignore first 5 characters, then match
julia> m = match(r"(\wat)", "cat, sat, mat", 5)
RegexMatch("sat", 1="sat")
在 Julia 中,match()只会找到目标字符串里的_第一个_匹配项:其他一些语言中那样的全局修饰符在这里并不存在。
取而代之的是 eachmatch(),它返回一个匹配项的迭代器。它的求值是惰性的,因此你可能需要把它转换成你想要的格式。
julia> matches = eachmatch(r"(\wat)", "cat, sat, mat")
Base.RegexMatchIterator{String}(r"(\wat)", "cat, sat, mat", false)
# convert to vector
julia> collect(matches)
3-element Vector{RegexMatch}:
RegexMatch("cat", 1="cat")
RegexMatch("sat", 1="sat")
RegexMatch("mat", 1="mat")
# convert with comprehension
julia> [m.match for m in matches]
3-element Vector{SubString{String}}:
"cat"
"sat"
"mat"
# broadcast an anonymous function
julia> (m -> m.match).(matches)
3-element Vector{SubString{String}}:
"cat"
"sat"
"mat"
默认不允许匹配项重叠。加上关键字参数overlap = true即可改变这一默认行为。
julia> eachmatch(r"aba", "abababa") |> collect # matches at positions 1, 5
2-element Vector{RegexMatch}:
RegexMatch("aba")
RegexMatch("aba")
julia> eachmatch(r"aba", "abababa"; overlap = true) |> collect # also matches at position 3
3-element Vector{RegexMatch}:
RegexMatch("aba")
RegexMatch("aba")
RegexMatch("aba")
使用正则表达式的一个常见原因,是把匹配项替换成另一个字符串。
replace()函数在 Strings概念中讨论过,当时用的是字符串字面量来搜索。同一个函数也可以充分发挥正则表达式匹配的全部威力。
julia> replace("some string", r"[aeiou]" => "*")
"s*m* str*ng"
julia> replace("first second", r"(\w+) (?<agroup>\w+)" => s"\g<agroup> \1")
"second first"
上面的第二个例子展示了,在s" "字符串中,编号捕获组和命名捕获组都可以用在替换里。
更多细节请参阅手册:这个话题总是让大多数程序员一次次翻回文档!