Python에서 str은 유니코드 코드 포인트로 이루어진 불변 시퀀스예요.
글자, 발음 구별 부호, 위치 지정 문자, 숫자, 통화 기호, 이모지, 문장 부호, 공백과 줄 바꿈 문자 등이 모두 포함될 수 있어요.
str은 불변이기 때문에 메모리에서 str 객체의 값은 바뀌지 않아요. 문자열을 수정하는 것처럼 보이는 메서드는 그 str 객체의 새 복사본이나 인스턴스를 반환해요.
str 리터럴은 작은따옴표 '나 큰따옴표 "로 선언할 수 있어요. 필요할 때는 이스케이프 문자 \를 사용할 수 있어요.
>>> single_quoted = 'These allow "double quoting" without "escape" characters.'
>>> double_quoted = "These allow embedded 'single quoting', so you don't have to use an 'escape' character."
>>> escapes = 'If needed, a \'slash\' can be used as an escape character within a string when switching quote styles won\'t work.'
여러 줄 문자열은 '''나 """로 선언해요.
>>> triple_quoted = '''Three single quotes or "double quotes" in a row allow for multi-line string literals.
Line break characters, tabs and other whitespace are fully supported.
You\'ll most often encounter these as "doc strings" or "doc tests" written just below the first line of a function or class definition.
They\'re often used with auto documentation ✍ tools.
'''
문자열은 + 연산자로 이어 붙일 수 있어요.
다만 이 방법은 성능이 좋지 않고 유지 보수하기도 쉽지 않으니 아껴서 쓰는 게 좋아요.
language = "Ukrainian"
number = "nine"
word = "дев'ять"
sentence = word + " " + "means" + " " + number + " in " + language + "."
>>> print(sentence)
...
"дев'ять means nine in Ukrainian."
list, tuple, set 같이 여러 개별 문자열로 이루어진 컬렉션을 하나의 str로 합쳐야 한다면 <str>.join(<iterable>)를 쓰는 편이 더 좋아요:
# str.join() makes a new string from the iterables elements.
>>> chickens = ["hen", "egg", "rooster"] # Lists are iterable.
>>> ' '.join(chickens)
'hen egg rooster'
# Any string can be used as the joining element.
>>> ' :: '.join(chickens)
'hen :: egg :: rooster'
>>> ' 🌿 '.join(chickens)
'hen 🌿 egg 🌿 rooster'
# Any iterable can be used as input.
>>> flowers = ("rose", "daisy", "carnation") # Tuples are iterable.
>>> '*-*'.join(flowers)
'rose*-*daisy*-*carnation'
>>> flowers = {"rose", "daisy", "carnation"} # Sets are iterable, but output order is not guaranteed.
>>> '*-*'.join(flowers)
'rose*-*carnation*-*daisy'
>>> phrase = "This is my string" # Strings are iterable, but be careful!
>>> '..'.join(phrase)
'T..h..i..s.. ..i..s.. ..m..y.. ..s..t..r..i..n..g'
# Separators are inserted **between** elements, but can be any string (including spaces).
# This can be exploited for interesting effects.
>>> under_words = ['under', 'current', 'sea', 'pin', 'dog', 'lay']
>>> separator = ' ⤴️ under'
>>> separator.join(under_words)
'under ⤴️ undercurrent ⤴️ undersea ⤴️ underpin ⤴️ underdog ⤴️ underlay'
# The separator can be composed different ways, as long as the result is a string.
>>> upper_words = ['upper', 'crust', 'case', 'classmen', 'most', 'cut']
>>> separator = ' 🌟 ' + upper_words[0]
>>> separator.join(upper_words)
'upper 🌟 uppercrust 🌟 uppercase 🌟 upperclassmen 🌟 uppermost 🌟 uppercut'
str 안의 코드 포인트는 왼쪽에서부터 0-based index 번호로 가리킬 수 있어요:
creative = '창의적인'
>>> creative[0]
'창'
>>> creative[2]
'적'
>>> creative[3]
'인'
인덱싱은 오른쪽에서도 할 수 있는데, 이때는 -1-based index부터 시작해요:
creative = '창의적인'
>>> creative[-4]
'창'
>>> creative[-2]
'적'
>>> creative[-1]
'인'
Python에는 별도의 "문자"나 "룬" 타입이 없어서, 문자열을 인덱싱하면 길이가 1인 새로운 str이 만들어져요:
>>> website = "exercism"
>>> type(website[0])
<class 'str'>
>>> len(website[0])
1
>>> website[0] == website[0:1] == 'e'
True
부분 문자열은 _슬라이스 표기법_으로 선택할 수 있어요. <str>[<start>:stop:<step>]를 사용하면 새로운 문자열이 만들어져요.
결과에는 stop 인덱스가 포함되지 않아요.
start를 지정하지 않으면 시작 인덱스는 0이 돼요.
stop을 지정하지 않으면 stop 인덱스는 문자열의 끝이 돼요.
moon_and_stars = '🌟🌟🌙🌟🌟⭐'
sun_and_moon = '🌞🌙🌞🌙🌞🌙🌞🌙🌞'
>>> moon_and_stars[1:4]
'🌟🌙🌟'
>>> moon_and_stars[:3]
'🌟🌟🌙'
>>> moon_and_stars[3:]
'🌟🌟⭐'
>>> moon_and_stars[:-1]
'🌟🌟🌙🌟🌟'
>>> moon_and_stars[:-3]
'🌟🌟🌙'
>>> sun_and_moon[::2]
'🌞🌞🌞🌞🌞'
>>> sun_and_moon[:-2:2]
'🌞🌞🌞🌞'
>>> sun_and_moon[1:-1:2]
'🌙🌙🌙🌙'
문자열은 <str>.split(<separator>)로 더 작은 문자열로 나눌 수도 있는데, 이때 부분 문자열의 list를 반환해요.
이 배열은 필요하다면 다시 인덱싱하거나 나눌 수 있어요.
<str>.split()을 인자 없이 사용하면 공백을 기준으로 문자열을 나눠요.
>>> cat_ipsum = "Destroy house in 5 seconds mock the hooman."
>>> cat_ipsum.split()
...
['Destroy', 'house', 'in', '5', 'seconds', 'mock', 'the', 'hooman.']
>>> cat_ipsum.split()[-1]
'hooman.'
>>> cat_words = "feline, four-footed, ferocious, furry"
>>> cat_words.split(', ')
...
['feline', 'four-footed', 'ferocious', 'furry']
<str>.split()의 구분자는 두 글자 이상일 수도 있어요.
나눌 때는 문자열 전체를 기준으로 일치 여부를 판단해요.
>>> colors = """red,
orange,
green,
purple,
yellow"""
>>> colors.split(',\n')
['red', 'orange', 'green', 'purple', 'yellow']
문자열은 모든 공통 시퀀스 연산을 지원해요.
개별 코드 포인트는 for item in <str>로 루프를 돌며 반복할 수 있어요.
인덱스와 항목을 함께 반복하려면 for index, item in enumerate(<str>)를 사용해요.
>>> exercise = 'လေ့ကျင့်'
# Note that there are more code points than perceived glyphs or characters
>>> for code_point in exercise:
... print(code_point)
...
လ
ေ
့
က
ျ
င
်
့
# Using enumerate will give both the value and index position of each element.
>>> for index, code_point in enumerate(exercise):
... print(index, ": ", code_point)
...
0 : လ
1 : ေ
2 : ့
3 : က
4 : ျ
5 : င
6 : ်
7 : ့
여동생의 영어 어휘 숙제를 도와주고 있어요. 여동생은 이 숙제를 아주 지루해하고 있어요. 여동생 반에서는 _접두사_와 _접미사_를 붙여서 새로운 단어를 만드는 법을 배우고 있어요. 선생님은 주어진 단어들에 접두사는 앞에, 접미사는 뒤에 붙여서 철자가 맞는, 올바르게 변형된 단어를 찾고 있어요.
이 과제에는 네 가지 활동이 있고, 각 활동마다 다룰 텍스트나 단어가 주어져요.
영어에서 가장 흔한 접두사 중 하나는 "아니다"라는 뜻의 un이에요.
이 활동에서 여동생은 단어에 un을 붙여서 부정적인, 즉 "아니다"라는 뜻의 단어를 만들어야 해요.
word를 매개변수로 받아서 un이 붙은 새로운 단어를 반환하는 add_prefix_un(<word>) 함수를 구현해요:
>>> add_prefix_un("happy")
'unhappy'
>>> add_prefix_un("manageable")
'unmanageable'
여동생 반에서 공부하고 있는 흔한 접두사가 네 개 더 있어요.
en('안에 넣다' 또는 '덮다'라는 뜻),
pre('앞' 또는 '앞으로'라는 뜻),
auto('스스로' 또는 '같은'이라는 뜻),
그리고 inter('사이' 또는 '가운데'라는 뜻)예요.
이 연습 문제에서는 반 친구들이 이 접두사들을 사용해서 어휘 단어 그룹을 만들고, 그래서 함께 공부할 수 있어요. 각 접두사는 함께 쓰이는 흔한 단어들과 함께 목록으로 주어져요. 학생들은 접두사를 적용해서, 그 접두사가 모든 단어에 적용된 모습을 보여 주는 문자열을 만들어야 해요.
다음과 같은 형태로 vocab_words를 매개변수로 받는 make_word_groups(<vocab_words>) 함수를 구현해요:
[<prefix>, <word_1>, <word_2> .... <word_n>].
그리고 각 단어에 접두사가 적용된, '<prefix> :: <prefix><word_1> :: <prefix><word_2> :: <prefix><word_n>'처럼 생긴 문자열을 반환해요.
입력을 처리하기 위해 for문이나 while문을 만들 필요는 없어요.
대신 어떤 문자열 메서드와 구분자를 사용할 수 있을지 잘 생각해 봐요.
>>> make_word_groups(['en', 'close', 'joy', 'lighten'])
'en :: enclose :: enjoy :: enlighten'
>>> make_word_groups(['pre', 'serve', 'dispose', 'position'])
'pre :: preserve :: predispose :: preposition'
>> make_word_groups(['auto', 'didactic', 'graph', 'mate'])
'auto :: autodidactic :: autograph :: automate'
>>> make_word_groups(['inter', 'twine', 'connected', 'dependent'])
'inter :: intertwine :: interconnected :: interdependent'
ness는 '~라는 상태'를 뜻하는 흔한 접미사예요.
이 활동에서 여동생은 ness 접미사를 떼어 내서 원래의 어근 단어를 찾아야 해요.
그런데 물론 성가신 철자 규칙이 있어요. 어근 단어가 원래 자음 뒤에 'y'로 끝났다면, 그 'y'는 'i'로 바뀌었어요.
'ness'를 떼어 낼 때는 그런 어근 단어의 'y'를 되살려야 해요. 예를 들어 happiness --> happi --> happy예요.
word를 받아서 ness 접미사가 없는 어근 단어를 반환하는 remove_suffix_ness(<word>) 함수를 구현해요.
>>> remove_suffix_ness("heaviness")
'heavy'
>>> remove_suffix_ness("sadness")
'sad'
접미사는 단어의 품사를 바꾸는 데 자주 쓰여요.
영어에서는 형용사에 en 접미사를 붙여 형용사가 동사가 되게 하는, 이른바 "동사화"가 흔한 관행이에요.
이 과제에서 여동생은 문장에서 형용사를 추출해서 동사로 바꾸는, 단어의 "동사화"를 연습할 거예요. 다행히 여기서 변형해야 하는 단어는 모두 "규칙" 단어라서, 접미사를 붙일 때 철자 변화가 필요 없어요.
매개변수 두 개를 받는 adjective_to_verb(<sentence>, <index>) 함수를 구현해요.
어휘 단어를 사용한 sentence와, 그 문장을 나눴을 때의 단어 index예요.
함수는 추출한 형용사를 동사로 반환해야 해요.
>>> adjective_to_verb('I need to make that bright.', -1 )
'brighten'
>>> adjective_to_verb('It got dark as the sun set.', 2)
'darken'
Exercism에 가입하고 Python 트랙을 개념 17개연습 문제 146개, 그리고 실제 사람의 멘토링과 함께 배우고 익혀 보세요. 모두 무료예요.