Tracce
/
jq
jq
/
Programma
/
Espressioni regolari
Es

Espressioni regolari in jq

1 esercizio

Informazioni su Espressioni regolari

Le espressioni regolari (regex) sono sequenze di caratteri che specificano un modello di ricerca in un testo.

Imparare la sintassi delle espressioni regolari va oltre lo scopo di questo argomento. Ci concentreremo sulle espressioni che jq fornisce per utilizzare le regex.

Variante di regex

Strumenti diversi implementano versioni diverse delle espressioni regolari. jq incorpora la libreria di regex Oniguruma, in gran parte compatibile con le regex di Perl v5.8.

La sintassi specifica usata da jq si può trovare nel repository GitHub di Oniguruma.

Caution

jq non ha alcuna sintassi speciale per le espressioni regolari. Sono semplicemente espresse come stringhe. Questo significa che eventuali backslash nell'espressione regolare devono essere escapati nella stringa.

Per esempio, la classe di caratteri per le cifre (\d) deve essere scritta come "\\d".

Funzioni regex

Le espressioni regolari in jq sono limitate a un insieme di filtri.

Corrispondenza semplice

Quando devi sapere se una stringa corrisponde a un modello, usa il filtro test.

STRING | test(REGEX)
STRING | test(REGEX; FLAGS)
STRING | test([REGEX, FLAGS])

Questo filtro restituisce un risultato booleano.

"Hello World!" | test("W")    # => true
"Goodbye Mars" | test("W")    # => false

Informazioni sulla corrispondenza

Quando devi estrarre la sottostringa che corrisponde effettivamente al modello, usa il filtro match.

STRING | match(REGEX)
STRING | match(REGEX; FLAGS)
STRING | match([REGEX, FLAGS])

Questo filtro restituisce:

  • niente se non c'è stata una corrispondenza, oppure
  • un oggetto che contiene varie proprietà se c'è stata una corrispondenza.

Questo esempio cerca due vocali consecutive identiche usando la sintassi di backreference, \1.

"Hello World!" | match("([aeiou])\\1")
# => empty

"Goodbye Mars" | match("([aeiou])\\1")
# => {
#      "offset": 1,
#      "length": 2,
#      "string": "oo",
#      "captures": [
#        {
#          "offset": 1,
#          "length": 1,
#          "string": "o",
#          "name": null
#        }
#      ]
#    }

Il filtro match restituisce un oggetto per ogni corrispondenza. Questo esempio mostra il flag "g" in azione per trovare tutte le vocali.

"Goodbye Mars" | match("[aeiou]"; "g")
# => { "offset": 1, "length": 1, "string": "o", "captures": [] }
#    { "offset": 2, "length": 1, "string": "o", "captures": [] }
#    { "offset": 6, "length": 1, "string": "e", "captures": [] }
#    { "offset": 9, "length": 1, "string": "a", "captures": [] }

Sottostringhe catturate

Simile al filtro match, il filtro capture restituisce un oggetto se c'è stata una corrispondenza.

STRING | capture(REGEX)
STRING | capture(REGEX; FLAGS)
STRING | capture([REGEX, FLAGS])

L'oggetto restituito è una mappatura delle catture con nome.

"JIRAISSUE-1234" | capture("(?<project>\\w+)-(?<issue_num>\\d+)")
# => {
#      "project": "JIRAISSUE",
#      "issue_num": "1234"
#    }

Solo le sottostringhe

Il filtro scan è simile a match con il flag "g".

STRING | scan(REGEX)
STRING | scan(REGEX; FLAGS)

# note, there is no scan([REGEX, FLAGS]) version, unlike other filters

scan produrrà un flusso di sottostringhe.

"Goodbye Mars" | scan("[aeiou]")
# => "o"
#    "o"
#    "e"
#    "a"

Usa il costruttore di array [...] per catturare le sottostringhe.

"Goodbye Mars" | [ scan("[aeiou]") ]
# => ["o", "o", "e", "a"]

Dividere una stringa

Se conosci le parti della stringa che vuoi mantenere, usa match o scan. Se conosci le parti che vuoi scartare, usa split.

STRING | split(REGEX; FLAGS)
Caution

Il filtro split a un solo argomento tratta il suo argomento come una stringa fissa.

Per usare una regex con split, devi fornire il secondo argomento; va bene usare una stringa vuota.

Un esempio che divide una stringa su spazi bianchi arbitrari.

"first   second           third fourth" | split("\\s+"; "")
# => ["first", "second", "third", "fourth"]
Note

Questo è quello che succede se dimentichiamo l'argomento dei flag.

"first   second           third fourth" | split("\\s+")
# => ["first   second           third fourth"]

Un solo risultato: la stringa fissa \s+ non è stata trovata nell'input.

Il filtro split a un argomento non può gestire spazi bianchi arbitrari. Dividere su uno spazio dà questo risultato.

"first   second           third fourth" | split(" ")
# => ["first", "", "", "second", "", "", "", "",
#     "", "", "", "", "", "", "third", "fourth" ]

Sostituzioni

I filtri sub e gsub possono trasformare la stringa di input, sostituendo le porzioni corrispondenti dell'input con una stringa di sostituzione.

Per sostituire solo la prima corrispondenza, usa sub. Per sostituire tutte le corrispondenze, usa gsub.

STRING | sub(REGEX; REPLACEMENT)
STRING | sub(REGEX; REPLACEMENT; FLAGS)
STRING | gsub(REGEX; REPLACEMENT)
STRING | gsub(REGEX; REPLACEMENT; FLAGS)
"Goodnight kittens. Goodnight mittens." | sub("night"; " morning")
# => "Good morning kittens. Goodnight mittens."

"Goodnight kittens. Goodnight mittens." | gsub("night"; " morning")
# => "Good morning kittens. Good morning mittens."

Il testo di sostituzione può fare riferimento alle sottostringhe corrispondenti; usa catture con nome e interpolazione di stringhe.

"Some 3-letter acronyms: gnu, csv, png"
| gsub( "\\b(?<tla>[[:alpha:]]{3})\\b";     # find words 3 letters long
        "\(.tla | ascii_upcase)" )          # upper-case the match
# => "Some 3-letter acronyms: GNU, CSV, PNG"

Flag

In tutti i filtri precedenti, FLAGS è una stringa composta da zero o più dei flag supportati.

  • g - Ricerca globale (trova tutte le corrispondenze, non solo la prima)
  • i - Ricerca senza distinzione tra maiuscole e minuscole
  • m - Modalità multi-riga ('.' corrisponderà ai caratteri di nuova riga)
  • n - Ignora le corrispondenze vuote
  • p - Sono abilitate sia la modalità s sia la modalità m
  • s - Modalità riga singola ('^' -> '\A', '$' -> '\Z')
  • l - Trova le corrispondenze più lunghe possibili
  • x - Formato regex esteso (ignora spazi bianchi e commenti)

Per esempio

"JIRAISSUE-1234" | capture("(?<project>\\w+)-(?<issue_num>\\d+)")

# or with Extended formatting

"JIRAISSUE-1234" | capture("
                     (?<project>   \\w+ )  # the Jira project
                     -                     # followed by a hyphen
                     (?<issue_num> \\d+ )  # followed by digits
                   "; "x")
Modifica tramite GitHub Il collegamento si apre in una nuova finestra o scheda

Impara Espressioni regolari

La pratica è bloccata

Sblocca un altro esercizio per esercitarti su Espressioni regolari