轨道
/
Python
Python
/
练习
/
宴会服务员
宴会服务员

宴会服务员

学习练习

简介

一个集合是由可哈希对象组成的_可变_且_无序_的容器。 集合的成员必须互不相同,不允许有重复项。 集合可以存放多种不同的数据类型,甚至像由tuple组成的tuples这样的嵌套结构,只要所有元素都能_被哈希_。 集合还有一种不可变的版本,叫做frozensets。

集合最常见的用途,是从其他数据结构或元素分组中快速去除重复项。 当不需要保持顺序、也不需要追踪重复项时,集合还常用于高效比较。

和其他容器类型(字典、数组、元组)一样,sets支持:

  • 通过for item in <set>进行迭代
  • 通过in和not in进行成员检查,
  • 通过len()计算长度,以及
  • 通过copy()进行浅拷贝

sets不支持:

  • 任何形式的下标访问
  • 通过排序或插入来控制顺序
  • 切片
  • 通过+拼接

在set中检查成员的时间复杂度是常数级(平均而言);而在list或string中检查成员时,时间复杂度会随着数据长度的增加而增长。 <set>.union()、<set>.intersection()或<set>.difference()等方法的时间复杂度同样是常数级(平均而言)。

集合字面量

可以直接把set写成_集合字面量_,用花括号{}把元素括起来,元素之间用逗号分隔。 重复项会被自动忽略:

>>> one_element = {'➕'}
{'➕'}

>>> multiple_elements = {'➕', '🔻', '🔹', '🔆'}
{'➕', '🔻', '🔹', '🔆'}

>>> multiple_duplicates =  {'Hello!', 'Hello!', 'Hello!', 
                            '¡Hola!','Привіт!', 'こんにちは!', 
                            '¡Hola!','Привіт!', 'こんにちは!'}
{'こんにちは!', '¡Hola!', 'Hello!', 'Привіт!'}

集合字面量和dict字面量使用相同的花括号,这意味着要创建空set,必须使用set()。

集合构造函数

set()(set类的构造函数)可以接受任何作为实参传入的iterable。 会逐个遍历iterable中的元素,并逐个添加到set中。 元素的顺序不会保留,重复项会被自动忽略:

# To create an empty set, the constructor must be used.
>>> no_elements = set()
set()

# The tuple is unpacked & each element is added.  
# Duplicates are removed.
>>> elements_from_tuple = set(("Parrot", "Bird", 
                               334782, "Bird", "Parrot"))
{334782, 'Bird', 'Parrot'}

# The list is unpacked & each element is added.
# Duplicates are removed.
>>> elements_from_list = set([2, 3, 2, 3, 3, 3, 5, 
                              7, 11, 7, 11, 13, 13])
{2, 3, 5, 7, 11, 13}

创建集合时的坑

由于“解包”行为,用set()处理字符串时,结果可能会出乎意料:

# String elements (Unicode code points) are 
# iterated through and added *individually*.
>>> elements_string = set("Timbuktu")
{'T', 'b', 'i', 'k', 'm', 't', 'u'}

# Unicode separators and positioning code points 
# are also added *individually*.
>>> multiple_code_points_string = set('अभ्यास')
{'अ', 'भ', 'य', 'स', 'ा', '्'}

集合可以存放不同的数据类型和_嵌套_的数据类型,但所有set元素都必须_可哈希_:

# Attempting to use a list for a set member throws a TypeError
>>> lists_as_elements = {['🌈','💦'], 
                        ['☁️','⭐️','🌍'], 
                        ['⛵️', '🚲', '🚀']}

Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
TypeError: unhashable type: 'list'


# Standard sets are mutable, so they cannot be hashed.
>>> sets_as_elements = {{'🌈','💦'}, 
                        {'☁️','⭐️','🌍'}, 
                        {'⛵️', '🚲', '🚀'}}

Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
TypeError: unhashable type: 'set'

使用集合

集合的方法通常模仿数学上的集合运算。 这些方法大多(但并非全部)都有对应的运算符。 方法通常接受任何iterable作为实参,而运算符则要求运算两侧都是sets或frozensets。

不相交集合

<set>.isdisjoint(<other_collection>)方法用于测试一个集合的元素是否与另一个set的元素有重叠。 这个方法接受任何iterable或set作为实参。 如果两个集合没有任何共同元素,就返回True;如果有共享元素,则返回False。 没有对应的运算符:

# Both mammals and additional_animals are lists.
>>> mammals = ['squirrel','dog','cat','cow', 'tiger', 'elephant']
>>> additional_animals = ['pangolin', 'panda', 'parrot', 
                          'lemur', 'tiger', 'pangolin']

# Animals is a dict.
>>> animals = {'chicken': 'white',
               'sparrow': 'grey',
               'eagle': 'brown and white',
               'albatross': 'grey and white',
               'crow': 'black',
               'elephant': 'grey', 
               'dog': 'rust',
               'cow': 'black and white',
               'tiger': 'orange and black',
               'cat': 'grey',
               'squirrel': 'black'}
               
# Birds is a set.
>>> birds = {'crow','sparrow','eagle','chicken', 'albatross'}

# Mammals and birds don't share any elements.
>>> birds.isdisjoint(mammals)
True

# There are also no shared elements between 
# additional_animals and birds.
>>> birds.isdisjoint(additional_animals)
True

# Animals and mammals have shared elements.
# **Note** The first object needs to be a set or converted to a set
# since .isdisjoint() is a set method.
>>> set(animals).isdisjoint(mammals)
False

子集与超集

<set>.issubset(<other_collection>)用于检查<set>中的每个元素是否也都在<other_collection>中。 对应的运算符形式是<set> <= <other_set>:

# Both mammals and additional_animals are lists.
>>> mammals = ['squirrel','dog','cat','cow', 'tiger', 'elephant']
>>> additional_animals = ['pangolin', 'panda', 'parrot', 
                          'lemur', 'tiger', 'pangolin']

# Animals is a dict.
>>> animals = {'chicken': 'white',
               'sparrow': 'grey',
               'eagle': 'brown and white',
               'albatross': 'grey and white',
               'crow': 'black',
               'elephant': 'grey', 
               'dog': 'rust',
               'cow': 'black and white',
               'tiger': 'orange and black',
               'cat': 'grey',
               'squirrel': 'black'}

# Birds is a set.
>>> birds = {'crow','sparrow','eagle','chicken', 'albatross'}

# Set methods will take any iterable as an argument.
# All members of birds are also members of animals.
>>> birds.issubset(animals)
True

# All members of mammals also appear in animals.
# **Note** The first object needs to be a set or converted to a set
# since .issubset() is a set method.
>>> set(mammals).issubset(animals)
True

# Both objects need to be sets to use a set operator
>>> birds <= set(mammals)
False

# A set is always a loose subset of itself.
>>> set(additional_animals) <= set(additional_animals)
True

<set>.issuperset(<other_collection>)是.issubset()的逆运算。 它用于检查<other_collection>中的每个元素是否也都在<set>中。 对应的运算符形式是<set> >= <other_set>:

# All members of mammals also appear in animals.
# **Note** The first object needs to be a set or converted to a set
# since .issuperset() is a set method.
>>> set(animals).issuperset(mammals)
True

# All members of animals do not show up as members of birds.
>>> birds.issuperset(animals)
False

# Both objects need to be sets to use a set operator
>>> birds >= set(mammals)
False

# A set is always a loose superset of itself.
>>> set(animals) >= set(animals)
True

集合的交集

<set>.intersection(*<other iterables>)返回一个新的set,其中的元素是原set和所有<others>共有的(换句话说,就是所有元素都相交的那个set)。 这个方法的运算符版本是<set> & <other set> & <other set 2> & ... <other set n>:

>>> perennials = {'Annatto','Asafetida','Asparagus','Azalea',
                 'Winter Savory', 'Broccoli','Curry Leaf','Fennel', 
                 'Kaffir Lime','Kale','Lavender','Mint','Oranges',
                 'Oregano', 'Tarragon', 'Wild Bergamot'}

>>> annuals = {'Corn', 'Zucchini', 'Sweet Peas', 'Marjoram', 
              'Summer Squash', 'Okra','Shallots', 'Basil', 
              'Cilantro', 'Cumin', 'Sunflower', 'Chervil', 
              'Summer Savory'}

>>> herbs = ['Annatto','Asafetida','Basil','Chervil','Cilantro',
            'Curry Leaf','Fennel','Kaffir Lime','Lavender',
            'Marjoram','Mint','Oregano','Summer Savory', 
            'Tarragon','Wild Bergamot','Wild Celery',
            'Winter Savory']


# Methods will take any iterable as an argument.
>>> perennial_herbs = perennials.intersection(herbs)
{'Annatto', 'Asafetida', 'Curry Leaf', 'Fennel', 'Kaffir Lime',
 'Lavender', 'Mint', 'Oregano', 'Wild Bergamot','Winter Savory'}

# Operators require both groups be sets.
>>> annuals & set(herbs)
 {'Basil', 'Chervil', 'Marjoram', 'Cilantro'}

集合的并集

<set>.union(*<other iterables>)返回一个新的set,其中的元素来自<set>和所有<other iterables>。 这个方法的运算符形式是<set> | <other set 1> | <other set 2> | ... | <other set n>:

>>> perennials = {'Asparagus', 'Broccoli', 'Sweet Potato', 'Kale'}
>>> annuals = {'Corn', 'Zucchini', 'Sweet Peas', 'Summer Squash'}
>>> more_perennials = ['Radicchio', 'Rhubarb', 
                      'Spinach', 'Watercress']

# Methods will take any iterable as an argument.
>>> perennials.union(more_perennials)
{'Asparagus','Broccoli','Kale','Radicchio','Rhubarb',
'Spinach','Sweet Potato','Watercress'}

# Operators require sets.
>>> set(more_perennials) | perennials
{'Asparagus',
 'Broccoli',
 'Kale',
 'Radicchio',
 'Rhubarb',
 'Spinach',
 'Sweet Potato',
 'Watercress'}

集合的差集

<set>.difference(*<other iterables>)返回一个新的set,其中的元素来自原<set>但不在<others>中。 这个方法的运算符版本是<set> - <other set 1> - <other set 2> - ...<other set n>。

>>> berries_and_veggies = {'Asparagus', 
                          'Broccoli', 
                          'Watercress', 
                          'Goji Berries', 
                          'Goose Berries', 
                          'Ramps', 
                          'Walking Onions', 
                          'Blackberries', 
                          'Strawberries', 
                          'Rhubarb', 
                          'Kale', 
                          'Artichokes', 
                          'Currants'}

>>> veggies = ('Asparagus', 'Broccoli', 'Watercress', 'Ramps',
               'Walking Onions', 'Rhubarb', 'Kale', 'Artichokes')

# Methods will take any iterable as an argument.
>>> berries = berries_and_veggies.difference(veggies)
{'Blackberries','Currants','Goji Berries',
 'Goose Berries', 'Strawberries'}

# Operators require sets.
>>> berries_and_veggies - berries
{'Artichokes','Asparagus','Broccoli','Kale',
'Ramps','Rhubarb','Walking Onions','Watercress'}

<set>.symmetric_difference(<other iterable>)返回一个新的set,其中包含存在于<set>或<other>中的元素,但不会同时存在于两者中。 这个方法的运算符版本是<set> ^ <other set>:

>>> plants_1 = {'🌲','🍈','🌵', '🥑','🌴', '🥭'}
>>> plants_2 = ('🌸','🌴', '🌺', '🌲', '🌻', '🌵')


# Methods will take any iterable as an argument.
>>> fruit_and_flowers = plants_1.symmetric_difference(plants_2)
>>> fruit_and_flowers
{'🌸', '🌺', '🍈', '🥑', '🥭','🌻' }


# Operators require both groups be sets.
>>> fruit_and_flowers ^ plants_1
{'🌲',  '🌸', '🌴', '🌵','🌺', '🌻'}

>>> fruit_and_flowers ^ set(plants_2)
{'🥭', '🌴', '🌵', '🍈', '🌲', '🥑'}
Note

对两个以上的集合做对称差运算,得到的set既会包含每个set独有的元素,也会包含该序列中两个以上set之间共享的元素(详情见维基百科上关于对称差的条目)。

要想只得到序列中每个sets独有的元素,需要单独一步把所有两两组合之间的交集汇总起来,再将其移除:

>>> one = {'black pepper','breadcrumbs','celeriac','chickpea flour',
           'flour','lemon','parsley','salt','soy sauce',
           'sunflower oil','water'}

>>> two = {'black pepper','cornstarch','garlic','ginger',
           'lemon juice','lemon zest','salt','soy sauce','sugar',
           'tofu','vegetable oil','vegetable stock','water'}

>>> three = {'black pepper','garlic','lemon juice','mixed herbs',
             'nutritional yeast', 'olive oil','salt','silken tofu',
             'smoked tofu','soy sauce','spaghetti','turmeric'}

>>> four = {'barley malt','bell pepper','cashews','flour',
            'fresh basil','garlic','garlic powder', 'honey',
            'mushrooms','nutritional yeast','olive oil','oregano',
            'red onion', 'red pepper flakes','rosemary','salt',
            'sugar','tomatoes','water','yeast'}

>>> intersections = (one & two | one & three | one & four | 
                     two & three | two & four | three & four)
...
{'black pepper','flour','garlic','lemon juice','nutritional yeast', 
'olive oil','salt','soy sauce', 'sugar','water'}

# The ^ operation will include some of the items in intersections, 
# which means it is not a "clean" symmetric difference - there
# are overlapping members.
>>> (one ^ two ^ three ^ four) & intersections
{'black pepper', 'garlic', 'soy sauce', 'water'}

# Overlapping members need to be removed in a separate step
# when there are more than two sets that need symmetric difference.
>>> (one ^ two ^ three ^ four) - intersections
...
{'barley malt','bell pepper','breadcrumbs', 'cashews','celeriac',
  'chickpea flour','cornstarch','fresh basil', 'garlic powder',
  'ginger','honey','lemon','lemon zest','mixed herbs','mushrooms',
  'oregano','parsley','red onion','red pepper flakes','rosemary',
  'silken tofu','smoked tofu','spaghetti','sunflower oil', 'tofu', 
  'tomatoes','turmeric','vegetable oil','vegetable stock','yeast'}

说明

你和你的商业伙伴经营着一家小型餐饮公司。你们刚刚答应为当地的一个烹饪俱乐部承办一场活动,主打“俱乐部最爱”菜品。这个俱乐部没什么承办大型活动的经验,需要在组织、采购、备餐和上菜方面得到帮助。你决定写几个小的 Python 脚本,加快整个筹备过程。

1. 清理菜品食材

活动要用的菜谱来自各种各样的渠道,其中的食材似乎有重复(甚至更多)的条目,你可不想最后买回一堆多余的东西! 在采购和烹饪开始之前,每道菜的食材清单都需要“清理”一遍。

实现 clean_ingredients(<dish_name>, <dish_ingredients>) 函数,它接收一个菜名和一个食材list。 这个函数应返回一个tuple,第一项是菜名,后面是去重后的食材set。

>>> clean_ingredients('Punjabi-Style Chole', ['onions', 'tomatoes', 'ginger paste', 'garlic paste', 'ginger paste', 'vegetable oil', 'bay leaves', 'cloves', 'cardamom', 'cilantro', 'peppercorns', 'cumin powder', 'chickpeas', 'coriander powder', 'red chili powder', 'ground turmeric', 'garam masala', 'chickpeas', 'ginger', 'cilantro'])

>>> ('Punjabi-Style Chole', {'garam masala', 'bay leaves', 'ground turmeric', 'ginger', 'garlic paste', 'peppercorns', 'ginger paste', 'red chili powder', 'cardamom', 'chickpeas', 'cumin powder', 'vegetable oil', 'tomatoes', 'coriander powder', 'onions', 'cilantro', 'cloves'})

2. 鸡尾酒和无酒精鸡尾酒

活动上既有鸡尾酒,也有“无酒精鸡尾酒”,也就是_不含_酒精的混合饮品。 你需要确保“无酒精鸡尾酒”确实不含酒精,而鸡尾酒确实_含_酒精。

实现 check_drinks(<drink_name>, <drink_ingredients>) 函数,它接收饮品名称和一个食材list。 如果饮品不含酒精成分,函数应返回饮品名称加上 "Mocktail";如果饮品含酒精,则返回饮品名称加上 "Cocktail"。 在本练习中,鸡尾酒只包含 sets_categories_data.py 里 ALCOHOLS 常量中的酒精成分:

>>> from sets_categories_data import ALCOHOLS 

>>> check_drinks('Honeydew Cucumber', ['honeydew', 'coconut water', 'mint leaves', 'lime juice', 'salt', 'english cucumber'])
...
'Honeydew Cucumber Mocktail'

>>> check_drinks('Shirley Tonic', ['cinnamon stick', 'scotch', 'whole cloves', 'ginger', 'pomegranate juice', 'sugar', 'club soda'])
...
'Shirley Tonic Cocktail'

3. 给菜品分类

宾客名单里有饮食需求各不相同的食客,你的员工需要把菜品分到 Vegan、Vegetarian、Paleo、Keto 和 Omnivore 这几类中。 只有当一道菜的所有食材都出现在某个类别的食材集合里时,这道菜才属于该类别。

实现 categorize_dish(<dish_name>, <dish_ingredients>) 函数,它接收一个菜名和这道菜食材的set。 函数应返回一个字符串,格式为 dish name: <CATEGORY>(即这道菜属于哪个餐食类别)。 给出的所有菜品都会“归入”从 sets_categories_data.py 导入的某个类别(VEGAN、VEGETARIAN、PALEO、KETO 或 OMNIVORE)。

>>> from sets_categories_data import VEGAN, VEGETARIAN, PALEO, KETO, OMNIVORE


>>> categorize_dish('Sticky Lemon Tofu', {'tofu', 'soy sauce', 'salt', 'black pepper', 'cornstarch', 'vegetable oil', 'garlic', 'ginger', 'water', 'vegetable stock', 'lemon juice', 'lemon zest', 'sugar'})
...
'Sticky Lemon Tofu: VEGAN'

>>> categorize_dish('Shrimp Bacon and Crispy Chickpea Tacos with Salsa de Guacamole', {'shrimp', 'bacon', 'avocado', 'chickpeas', 'fresh tortillas', 'sea salt', 'guajillo chile', 'slivered almonds', 'olive oil', 'butter', 'black pepper', 'garlic', 'onion'})
...
'Shrimp Bacon and Crispy Chickpea Tacos with Salsa de Guacamole: OMNIVORE'

4. 标注过敏原和限制性食材

有些客人有过敏症和额外的饮食限制。 这些食材需要为每道菜做标记或注明,以免引起麻烦。

实现 tag_special_ingredients(<dish>) 函数,它接收一个tuple,第一个位置是菜名,第二个位置是这道菜的食材list或set。 返回菜名,后面跟上需要在菜品描述中特别注明的食材set。 list里的食材可能有重复,也可能没有。 在本练习中,所有需要标注的过敏原或特殊食材都在从 sets_categories_data.py 导入的 SPECIAL_INGREDIENTS 常量里。

>>> from sets_categories_data import SPECIAL_INGREDIENTS

>>> tag_special_ingredients(('Ginger Glazed Tofu Cutlets', ['tofu', 'soy sauce', 'ginger', 'corn starch', 'garlic', 'brown sugar', 'sesame seeds', 'lemon juice']))
...
('Ginger Glazed Tofu Cutlets', {'garlic','soy sauce','tofu'})

>>> tag_special_ingredients(('Arugula and Roasted Pork Salad', ['pork tenderloin', 'arugula', 'pears', 'blue cheese', 'pine nuts', 'balsamic vinegar', 'onions', 'black pepper']))
...
('Arugula and Roasted Pork Salad', {'pork tenderloin', 'blue cheese', 'pine nuts', 'onions'})

5. 汇总食材“总清单”

为了准备下单和采购,你需要为菜单上的所有东西汇总一份食材“总清单”(数量稍后再补)。

实现 compile_ingredients(<dishes>) 函数,它接收一个菜品list,返回所有列出菜品中全部食材的集合。 每道菜用它自己的食材set表示。

dishes = [ {'tofu', 'soy sauce', 'ginger', 'corn starch', 'garlic', 'brown sugar', 'sesame seeds', 'lemon juice'},
           {'pork tenderloin', 'arugula', 'pears', 'blue cheese', 'pine nuts',
           'balsamic vinegar', 'onions', 'black pepper'},
           {'honeydew', 'coconut water', 'mint leaves', 'lime juice', 'salt', 'english cucumber'}]

>>> compile_ingredients(dishes)
...
{'arugula', 'brown sugar', 'honeydew', 'coconut water', 'english cucumber', 'balsamic vinegar', 'mint leaves', 'pears', 'pork tenderloin', 'ginger', 'blue cheese', 'soy sauce', 'sesame seeds', 'black pepper', 'garlic', 'lime juice', 'corn starch', 'pine nuts', 'lemon juice', 'onions', 'salt', 'tofu'}

6. 挑出装盘传递的前菜

主人给了你一份菜品清单,希望把它们做成“一口大小”的前菜,装在托盘里上桌。 你需要把这些菜从按大份量准备的主菜品清单里挑出来。

实现 separate_appetizers(<dishes>, <appetizers>) 函数,它接收菜品名称的list和前菜名称的list。 函数应返回去掉前菜名称后的菜品名称list。 <dishes>或<appetizers>这两个list都可能含有重复项,需要去重。

dishes =    ['Avocado Deviled Eggs','Flank Steak with Chimichurri and Asparagus', 'Kingfish Lettuce Cups',
             'Grilled Flank Steak with Caesar Salad','Vegetarian Khoresh Bademjan','Avocado Deviled Eggs',
             'Barley Risotto','Kingfish Lettuce Cups']
          
appetizers = ['Kingfish Lettuce Cups','Avocado Deviled Eggs','Satay Steak Skewers',
              'Dahi Puri with Black Chickpeas','Avocado Deviled Eggs','Asparagus Puffs',
              'Asparagus Puffs']
              
>>> separate_appetizers(dishes, appetizers)
...
['Vegetarian Khoresh Bademjan', 'Barley Risotto', 'Flank Steak with Chimichurri and Asparagus', 
 'Grilled Flank Steak with Caesar Salad']

7. 找出只在一道菜谱里用到的食材

在每个类别(Vegan、Vegetarian、Paleo、Keto、Omnivore)里,你都要挑出只在一道菜中出现的食材。 这些“单例”食材会分配给专门的采购员,确保在赶着处理其他事情时不会把它们忘掉。

实现 singleton_ingredients(<dishes>, <INTERSECTIONS>) 函数,它接收一个菜品list和同一类别的 <CATEGORY>_INTERSECTIONS 常量。 每道菜用它的食材set表示。 每个 <CATEGORY>_INTERSECTIONS 都是一个set,包含在该类别中出现在不止一道菜里的食材。 利用集合运算,函数应返回“单例”食材的set(即在该类别中只出现在一道菜里的食材)。

from sets_categories_data import example_dishes, EXAMPLE_INTERSECTION

>>> singleton_ingredients(example_dishes, EXAMPLE_INTERSECTION)
...
{'garlic powder', 'sunflower oil', 'mixed herbs', 'cornstarch', 'celeriac', 'honey', 'mushrooms', 'bell pepper', 'rosemary', 'parsley', 'lemon', 'yeast', 'vegetable oil', 'vegetable stock', 'silken tofu', 'tofu', 'cashews', 'lemon zest', 'smoked tofu', 'spaghetti', 'ginger', 'breadcrumbs', 'tomatoes', 'barley malt', 'red pepper flakes', 'oregano', 'red onion', 'fresh basil'}
通过 GitHub 编辑 链接将在新窗口或新标签页中打开
Python Exercism

准备好开始 宴会服务员 了吗?

注册 Exercism,借助 17 个概念146 个练习 和真人导师指导,学习并掌握 Python,全部免费。