정규 표현식은 문자열을 패턴 매칭 하는 매우 다재다능한 방법이에요. 이를 위해 설계된 도메인 특화 언어(DSL)를 사용해요.
다른 여러 프로그래밍 언어와 마찬가지로, Julia는 자체 정규식 라이브러리를 구현하려고 하지 않아요. 대신 널리 쓰이는 PCRE2 라이브러리를 감싸서, (예를 들어) Gleam과 동일하고 JavaScript와 아주 비슷한 정규식 문법을 제공해요.
이 Julia 실러버스는 이미 기본적인 정규식 문법에 익숙하다고 가정해요. 여기서는 Julia에 특화된 기능에만 집중할게요.
정규 표현식 지식을 되살리는 데 도움이 될 자료 몇 가지를 아래에 정리했어요.
Julia의 정규 표현식 인터페이스는 매뉴얼에 설명되어 있어요.
Julia에서 정규 표현식은 여는 " 앞에 r을 붙인 문자열일 뿐이에요. 기본적인 기능은 모두 표준 라이브러리에 들어 있어요.
사실 Strings 개념에서 이미 다룬 함수 중 상당수는 occursin()처럼 기본적으로 정규식 검색을 위해 설계되어 있어요.
julia> re = r"test$"
r"test$"
julia> typeof(re)
Regex
# Does a string end with "test"?
julia> occursin(re, "this is a test")
true
julia> occursin(re, "these are tests")
false
닫는 따옴표 뒤에는 수정자 문자를 붙일 수 있어요. 예를 들어 i는 대소문자를 구분하지 않는 매치를 뜻해요.
julia> occursin(r"test", "Testing")
false
julia> occursin(r"test"i, "Testing")
true
흔히 우리는 무엇이 매치되는지 알고 싶어 해요. 이럴 때는 정규식 안의 괄호에 캡처 그룹을 넣고 match() 함수를 사용하면 돼요.
julia> m = match(r"(\d+g) .* (\d+ml)", "dissolve 25g sugar in 200ml water")
RegexMatch("25g sugar in 200ml", 1="25g", 2="200ml")
julia> m.captures
2-element Vector{Union{Nothing, SubString{String}}}:
"25g"
"200ml"
# how many matches?
julia> length(m.captures)
2
# what matched?
julia> m[1], m[2]
("25g", "200ml")
# Starting positions of the matches (character index)
julia> m.offsets
2-element Vector{Int64}:
10
23
물론 매치는 실패할 수도 있어요. 그러면 결과는 RegexMatch 대신 특별한 값 Nothing이 되니, 이 경우를 검사할 준비를 해 둬요.
# failed match
m = match(r"(not here)", "dissolve 25g sugar in 200ml water")
julia> typeof(m)
Nothing
julia> isnothing(m)
true
match는 기본적으로 문자열의 맨 앞에서 시작하지만, 처음 n개 문자를 무시하도록 오프셋 n을 지정할 수도 있어요.
# capture first match
julia> m = match(r"(\wat)", "cat, sat, mat")
RegexMatch("cat", 1="cat")
# ignore first 5 characters, then match
julia> m = match(r"(\wat)", "cat, sat, mat", 5)
RegexMatch("sat", 1="sat")
Julia에서 match()는 대상 문자열 안에서 첫 번째 매치만 찾아요. 일부 다른 언어에 있는 전역 수정자는 없어요.
대신 eachmatch()가 있어요. 이 함수는 매치의 이터레이터를 반환해요. 지연 평가되므로 원하는 형식으로 변환해야 할 수도 있어요.
julia> matches = eachmatch(r"(\wat)", "cat, sat, mat")
Base.RegexMatchIterator{String}(r"(\wat)", "cat, sat, mat", false)
# convert to vector
julia> collect(matches)
3-element Vector{RegexMatch}:
RegexMatch("cat", 1="cat")
RegexMatch("sat", 1="sat")
RegexMatch("mat", 1="mat")
# convert with comprehension
julia> [m.match for m in matches]
3-element Vector{SubString{String}}:
"cat"
"sat"
"mat"
# broadcast an anonymous function
julia> (m -> m.match).(matches)
3-element Vector{SubString{String}}:
"cat"
"sat"
"mat"
겹치는 매치는 기본적으로 허용되지 않아요. 이걸 허용하려면 overlap = true를 키워드 인자로 넘겨요.
julia> eachmatch(r"aba", "abababa") |> collect # matches at positions 1, 5
2-element Vector{RegexMatch}:
RegexMatch("aba")
RegexMatch("aba")
julia> eachmatch(r"aba", "abababa"; overlap = true) |> collect # also matches at position 3
3-element Vector{RegexMatch}:
RegexMatch("aba")
RegexMatch("aba")
RegexMatch("aba")
정규식을 쓰는 흔한 이유 중 하나는 매치된 부분을 다른 문자열로 바꾸기 위해서예요.
replace() 함수는 Strings 개념에서 문자열 리터럴로 검색하는 용도로 다뤘어요. 같은 함수를 쓰면 정규식 매치의 강력함을 온전히 활용할 수 있어요.
julia> replace("some string", r"[aeiou]" => "*")
"s*m* str*ng"
julia> replace("first second", r"(\w+) (?<agroup>\w+)" => s"\g<agroup> \1")
"second first"
위의 두 번째 예시는 s" " 문자열 안에서 번호가 붙은 캡처 그룹과 이름이 붙은 캡처 그룹을 모두 바꾸는 데 사용할 수 있다는 걸 보여줘요.
자세한 내용은 매뉴얼을 참고해요. 이 주제는 대부분의 프로그래머를 끊임없이 문서로 되돌아가게 만드는 주제예요!