正規表示式是一種用途極廣的方式,可用來對字串進行_模式比對_,它使用的是一套專為此目的設計的領域特定語言(DSL)。
和許多其他程式語言一樣,Julia 並不打算自己實作 Regex 函式庫,而是包裝廣受使用的 PCRE2 函式庫,因此提供與 Gleam(舉例來說)相同的 Regex 語法,而與 Javascript 非常相似。
這份 Julia 課程大綱假設你已經熟悉基本的 Regex 語法。我們只會專注在 Julia 特有的功能。
以下列出一些能幫你複習正規表示式知識的資源。
Julia 的正規表示式介面在手冊中有詳細說明。
在 Julia 裡,正規表示式就只是一個字串,在開頭的"前面加上r而已。所有基本功能都屬於標準函式庫的一部分。
事實上,在Strings概念中介紹過的許多函式,標準用法就是為了 Regex 搜尋而設計的,例如occursin()。
julia> re = r"test$"
r"test$"
julia> typeof(re)
Regex
# Does a string end with "test"?
julia> occursin(re, "this is a test")
true
julia> occursin(re, "these are tests")
false
修飾字元可以接在結尾的引號後面,例如用i表示不區分大小寫的比對。
julia> occursin(r"test", "Testing")
false
julia> occursin(r"test"i, "Testing")
true
我們通常會想知道到底比對到_什麼_。做法是在正規表示式中用括號加入擷取群組,然後使用match()函式。
julia> m = match(r"(\d+g) .* (\d+ml)", "dissolve 25g sugar in 200ml water")
RegexMatch("25g sugar in 200ml", 1="25g", 2="200ml")
julia> m.captures
2-element Vector{Union{Nothing, SubString{String}}}:
"25g"
"200ml"
# how many matches?
julia> length(m.captures)
2
# what matched?
julia> m[1], m[2]
("25g", "200ml")
# Starting positions of the matches (character index)
julia> m.offsets
2-element Vector{Int64}:
10
23
當然,比對也可能失敗。這時結果會是特殊值Nothing,而不是RegexMatch,所以要準備好檢查這種情況。
# failed match
m = match(r"(not here)", "dissolve 25g sugar in 200ml water")
julia> typeof(m)
Nothing
julia> isnothing(m)
true
雖然match預設從字串開頭開始,我們也可以指定位移n,忽略前n個字元。
# capture first match
julia> m = match(r"(\wat)", "cat, sat, mat")
RegexMatch("cat", 1="cat")
# ignore first 5 characters, then match
julia> m = match(r"(\wat)", "cat, sat, mat", 5)
RegexMatch("sat", 1="sat")
在 Julia 中,match()只會找到目標字串裡的_第一個_比對結果,並不像某些其他語言那樣有全域修飾字元。
取而代之的是eachmatch(),它會回傳一個比對結果的迭代器。由於它是惰性求值,你可能需要將它轉成你想要的格式。
julia> matches = eachmatch(r"(\wat)", "cat, sat, mat")
Base.RegexMatchIterator{String}(r"(\wat)", "cat, sat, mat", false)
# convert to vector
julia> collect(matches)
3-element Vector{RegexMatch}:
RegexMatch("cat", 1="cat")
RegexMatch("sat", 1="sat")
RegexMatch("mat", 1="mat")
# convert with comprehension
julia> [m.match for m in matches]
3-element Vector{SubString{String}}:
"cat"
"sat"
"mat"
# broadcast an anonymous function
julia> (m -> m.match).(matches)
3-element Vector{SubString{String}}:
"cat"
"sat"
"mat"
預設不允許重疊的比對結果。加上關鍵字引數overlap = true即可覆寫這個行為。
julia> eachmatch(r"aba", "abababa") |> collect # matches at positions 1, 5
2-element Vector{RegexMatch}:
RegexMatch("aba")
RegexMatch("aba")
julia> eachmatch(r"aba", "abababa"; overlap = true) |> collect # also matches at position 3
3-element Vector{RegexMatch}:
RegexMatch("aba")
RegexMatch("aba")
RegexMatch("aba")
使用 Regex 的一個常見理由,就是把比對到的內容取代成另一個字串。
replace()函式在Strings概念中已經介紹過,當時是用字串常值來搜尋。同一個函式也能充分發揮 Regex 比對的威力。
julia> replace("some string", r"[aeiou]" => "*")
"s*m* str*ng"
julia> replace("first second", r"(\w+) (?<agroup>\w+)" => s"\g<agroup> \1")
"second first"
上面的第二個例子展示了,在s" "字串中,編號與具名的擷取群組都能用在取代字串裡。
更多細節請參閱手冊:這個主題總是讓大多數程式設計師不斷回頭翻文件!