正規表示式(regex)是在 Elixir 中處理字串的強大工具。Elixir 的正規表示式遵循 PCRE 規範(Perl Compatible Regular Expressions)。代表正規表示式意義的字串模式會先經過編譯,再用來比對字串的全部或部分內容。
在 Elixir 中,建立正規表示式最常見的方式是使用~r sigil。Sigil 為 Elixir 中常見的工作提供了_語法糖_的捷徑。要對_字串字面值_進行比對,我們可以把字串本身接在 sigil 之後當作模式。
~r/test/
=~/2 運算子可用來對字串執行正規表示式比對,並回傳 boolean 結果。
"this is a test" =~ ~r/test/
# => true
使用 sigil 時有兩點要注意:
/
\)使用方括號 [] 比對字元範圍會定義一個_字元類別_。它會將任一個字元比對到類別中的字元。你也可以指定像 a-z 這樣的字元範圍,只要頭尾代表一段連續的碼點即可。
regex = ~r/[a-z][ADKZ][0-9][!?]/
"jZ5!" =~ regex
# => true
"jB5?" =~ regex
# => false
_簡寫字元類別_能讓模式更簡潔。例如:
\d 是 [0-9] 的簡寫(任何數字)\w 是 [A-Za-z0-9_] 的簡寫(任何「詞」字元)\s 是 [ \t\r\n\f] 的簡寫(任何空白字元)當_簡寫字元類別_用在 sigil 之外時,就必須跳脫:"\\d"
選擇_使用 | 作為特殊字元,表示比對其中_一個_或_另一個
regex = ~r/cat|bat/
"bat" =~ regex
# => true
"cat" =~ regex
# => true
_量詞_允許正規表示式中出現重複的模式。它們會影響量詞前面的群組。
{N, M},其中 N 是最少重複次數,M 是最多重複次數{N,} 比對 N 次以上的重複
{0,} 也可以寫成 *:比對零次以上的重複{1,} 也可以寫成 +:比對一次以上的重複{,N} 比對最多 N 次的重複圓括號 () 用來表示_群組_和_擷取_。在某些情況下,群組也可以被_擷取_並回傳以供使用。在 Elixir 中,這些可以是具名的或未具名的。在開括號之後加上 ?<name> 即可為擷取命名。群組會像單一單位般運作,例如後面接著_量詞_時。
regex = ~r/(h)at/
Regex.replace(regex, "hat", "\\1op")
# => "hop"
regex = ~r/(?<letter_b>b)/
Regex.scan(regex, "blueberry", capture: :all_names)
# => [["b"], ["b"]]
_錨點_用來將正規表示式綁定到要比對字串的開頭或結尾:
^ 錨定到字串的開頭$ 錨定到字串的結尾由於 ~r 是 "pattern" |> Regex.escape() |> Regex.compile!() 的捷徑,你也可以使用字串內插來動態建立正規表示式的模式:
anchor = "$"
regex = ~r/end of the line#{anchor}/
"end of the line?" =~ regex
# => false
"end of the line" =~ regex
# => true
你的任務是撰寫一個會接收事件的服務。每個事件都帶有一個日期,但你注意到有三種不同的格式被送到你服務的端點:
"01/01/1970""January 1, 1970""Thursday, January 1, 1970"你可以看出它們之間有一些相似之處,於是決定撰寫一些可組合的正規表達式模式。
實作 day/0、month/0 和 year/0,讓它們回傳一個字串模式,這個模式在編譯後會比對 "01/01/1970"(dd/mm/yyyy)中的數字部分。日和月可能以 1 或 01 的形式出現(左側補零)。
"31" =~ DateParser.day() |> Regex.compile!()
# => true
"12" =~ DateParser.month() |> Regex.compile!()
# => true
"1970" =~ DateParser.year() |> Regex.compile!()
# => true
實作 day_names/0 和 month_names/0,讓它們回傳一個字串模式,這個模式在編譯後會分別比對以名稱表示的星期幾和月份。
"Tuesday" =~ DateParser.day_names() |> Regex.compile!()
# => true
"June" =~ DateParser.month_names() |> Regex.compile!()
# => true
實作 capture_day/0、capture_month/0、capture_year/0、capture_day_name/0、capture_month_name/0,讓它們回傳一個字串模式,這個模式會把對應的部分捕獲到 "day"、"month"、"year"、"day_name"、"month_name" 這些名稱下。
DateParser.capture_month_name()
|> Regex.compile!()
|> Regex.named_captures("December")
# => %{"month_name" => "December"}
實作 capture_numeric_date/0、capture_month_name_date() 和 capture_day_month_name_date/0,讓它們回傳一個字串模式,這個模式會依照對應的日期格式,捕獲第 3 部分中的各個部分:
"01/01/1970"
"January 1, 1970"
"Thursday, January 1, 1970"
DateParser.capture_numeric_date()
|> Regex.compile!()
|> Regex.named_captures("01/01/1970")
# => %{"day" => "01", "month" => "01", "year" => "1970"}
實作 match_numeric_date/0、match_month_name_date/0 和 match_day_month_name_date/0,讓它們回傳一個已編譯的正規表達式,這個表達式只會比對日期,同時也能捕獲各個部分。
"Thursday, January 1, 1970 was the Unix epoch." =~ DateParser.match_day_month_name_date()
# => false
"Thursday, January 1, 1970" =~ DateParser.match_day_month_name_date()
# => true
DateParser.match_day_month_name_date()
|> Regex.named_captures("Thursday, January 1, 1970")
# => %{
# "day" => "1",
# "day_name" => "Thursday",
# "month_name" => "January",
# "year" => "1970"
# }