Add a fourth module,
a_package/module5.py, with its ownfunction5()and a top-levelprint()so the module announces itself when it loads. Import it three ways, usingimport a_package.module5,from a_package import module5, andfrom a_package.module5 import function5, and confirm the loading message prints once, however many of the three you use together.
Packages
shows the three import forms and the messages each module prints
as it loads. Python records every loaded module in
sys.modules under its full dotted name. Ask what a
second import of a name in that cache does with the
module’s top-level code.
# The shape of a_package/module5.py
def function5():
...The package’s __init__.py is the chapter’s:
# a_package/__init__.py
print("initializing a_package")# a_package/module5.py
print("importing module5 in a_package")
def function5():
return "function5 in module5 in a_package"# use_module5.py
import a_package.module5
from a_package import module5
from a_package.module5 import function5
#: initializing a_package
#: importing module5 in a_package
print(a_package.module5.function5())
#: function5 in module5 in a_package
print(module5.function5())
#: function5 in module5 in a_package
print(function5())
#: function5 in module5 in a_packageThe "importing module5..." message prints once,
no matter how many of the three import styles you combine.
Python caches every module in sys.modules on its
first import, keyed by the module’s full dotted name. A later
import of the same module, in any of these forms,
finds the cached module and skips running its top-level code
again. It binds a name to the module in the cache. The package’s
__init__.py runs once for the same reason.
Add
a_package/b_package/module6.pywith afunction6()that callsfunction5()frommodule5. Import and callfunction6()from a script outsidea_package, then renameb_packagetobPackage(and rename it back afterward) and explain, from the rules in File Names, why that name is a poor choice although the import still works.
Imports
Within a Package covers a module that imports a sibling, and
Nested
Packages shows the two-level import. module6
reaches module5 across a package boundary, so write
the import with its full a_package. path. For the
rename, compare bPackage with the naming advice in
File
Names.
# The shape of a_package/b_package/module6.py
from a_package.module5 import function5
def function6():
...If you write the import as
from module5 import function5, dropping the
a_package. prefix, both packages initialize and
then use_module6.py stops with
ModuleNotFoundError: No module named 'module5'.
Python looks for a top-level module5 on
sys.path, and no directory there holds a
module5.py, because the file lives inside
a_package. The full dotted path names
module5 through the package that contains it, so
the solution writes
from a_package.module5 import function5.
b_package keeps the chapter’s
__init__.py too:
# a_package/b_package/__init__.py
print("initializing b_package")# a_package/b_package/module6.py
from a_package.module5 import function5
print("importing module6 in b_package")
def function6():
return f"function6 calls {function5()}"# use_module6.py
from a_package.b_package.module6 import function6
#: initializing a_package
#: initializing b_package
#: importing module5 in a_package
#: importing module6 in b_package
print(function6())
#: function6 calls function5 in module5 in a_packageReach the parent package’s module. The
import crosses a package boundary, from b_package
up to a_package, so the absolute form is the right
choice here. The relative equivalent,
from ..module5 import function5, works too. Prefer
the relative form only for siblings within one package.
Load the import chain first. Python
initializes both packages before it runs module6.
module5 then loads before module6
finishes loading, because module6’s own import runs
while its body is executing. All four messages therefore print
before the script’s own print() runs.
After you rename the directory to bPackage and
update the import to a_package.bPackage.module6,
the script still runs. Python accepts any valid identifier as a
package name. The rename gives up what the convention provides.
File
Names calls for short, all-lowercase package names, so
bPackage stands out as something other than a
package to anyone scanning an import line. Its capital letter
also makes the name easy to mistype: on a case-insensitive
filesystem the shell and the editor accept bpackage
as well, and Python’s case-sensitive import check, the one
exercise 4 examines, rejects bpackage.
Write a small module
noisy2.pywhose top-level body prints a message, likenoisy.py. In a new script,lazy importbothnoisyandnoisy2, then usenoisy2beforenoisy. Confirm the two loading messages print in the order you used the modules, not the order you wrote thelazy importlines.
Watching
the Deferral runs a lazy import of a module
that prints when it loads. A lazy import binds the
name without running the module. The body runs when the name is
first used, so order your uses and your print()
calls to watch the messages appear.
# The shape of noisy.py
def announce():
...# The shape of noisy2.py
def announce():
...# noisy.py
print("noisy module loaded")
def announce():
print("noisy.announce() called")# noisy2.py
print("noisy2 module loaded")
def announce():
print("noisy2.announce() called")# lazy_demo.py
lazy import noisy
lazy import noisy2
print("before any use")
#: before any use
noisy2.announce()
#: noisy2 module loaded
#: noisy2.announce() called
print("between")
#: between
noisy.announce()
#: noisy module loaded
#: noisy.announce() called
print("after both")
#: after bothAlthough the lazy import noisy line comes first,
noisy’s body does not run until
noisy.announce() executes, and that call comes
after noisy2.announce(). Each
lazy import reserves the name. The module’s
top-level code runs at the first use of that name, so use order,
not declaration order, decides which module loads first.
module.py to
Module.pyRename
module.pytoModule.py, changeuse_module.pytoimport Module, and update its call toModule.useful_function(). Run it. Then change the import back toimport module, leaving the file namedModule.py, and run it again. Predict the result before you run it, then explain what you see, given that Windows and macOS openmodule.pyandModule.pyas the same file. Look upPYTHONCASEOKto confirm your explanation.
File
Names explains how Python compares an import name with the
file name on disk. The filesystem may treat the two names as one
file, but the import system reads the directory listing and
decides for itself. The PYTHONCASEOK environment
variable switches that comparison off, which tests your
explanation.
The import Module statement resolves, because
the name and the file agree, and the call in the body becomes
Module.useful_function() to match; left as
module.useful_function(), the call raises a
NameError. Changing the import back to
import module while the file is still
Module.py raises
ModuleNotFoundError: No module named 'module'. Did you mean: 'Module'?,
and it does so on every platform, Windows and macOS included.
The suggestion shows that Python found the file and declined
it.
The failure on Windows and macOS is the surprising part.
Windows’s NTFS and macOS’s default filesystem both open module.py and
Module.py as the same file, so the filesystem would
hand Python the file under either name. Python declines it. Its
import machinery reads the directory listing and compares the
module name against the name on disk case-sensitively, so
"module" + ".py" does not match the stored
Module.py and the search moves on.
The case check is deliberate, and PEP 235 says why: without it, a program written on Windows imports happily there and fails the first time it runs on Linux, where the two names really are different files. Making the case rule the same everywhere turns a portability bug that surfaces in someone else’s CI into one that surfaces on your own machine.
Setting PYTHONCASEOK in the environment turns
the check off on a case-insensitive platform, and
import module then finds Module.py.
The variable exists for legacy code. Leave it unset in new code.
A Python environment variable turns the check off, so the check
is Python’s rather than the filesystem’s.
None of this arises if you follow the convention. File
Names recommends snake_case for modules, and
with an all-lowercase name, an import and its file cannot differ
in case.
Change
a_package/module4.pyto the absolute importfrom a_package.module1 import function1and confirmuse_module4.pystill works. Then runpython a_package/module4.pydirectly, both before and after the change. Both fail, with different errors: explain each, and say whypython -m a_package.module4works either way.
Imports
Within a Package shows module4’s relative
import, and PYTHONPATH
describes where Python searches for top-level names. A relative
import needs the module’s __package__, and a file
run as a script has none. An absolute import needs the project
root on sys.path, which depends on the source of
the first entry. Compare what python file.py and
python -m package.module put there.
Changing a_package/module4.py to
from a_package.module1 import function1 leaves use_module4.py working as
before. Both forms find the same function. They differ only in
how they name it.
Running the module directly fails either way, with different
errors. With the relative import,
python a_package/module4.py reports:
ImportError: attempted relative import with no known parent package
Python resolves a relative import against the module’s
__package__, and a file run as a script has none:
it runs as __main__, which belongs to no package.
The single dot has no parent to name.
With the absolute import, the same command reports:
ModuleNotFoundError: No module named 'a_package'. Did you mean: 'b_package'?
The name is now fully qualified, so the parent question does
not arise. But sys.path[0] is the directory of the
script you ran, a_package/. The project root is
nowhere on the path, so the search for a top-level package
called a_package fails: Python is inside the
package, looking for it. The suggestion names
b_package, the one package Python does find on that
path.
python -m a_package.module4 works with either
form, and fixes both problems at once. -m sets
sys.path[0] to the current directory rather than
the script’s, so a_package is findable. It also
imports the module as a member of its package rather than
running a loose file, so __package__ holds
a_package and the dot resolves.
A module inside a package is not a script. -m is
how you run a package module, and a file you intend to run both
ways belongs at the top level, outside any package.
__all__Remove the
__all__line fromexporting.py. Predict whatstar_import.pyprints without it, run it to check, then restore the line.
What
a Module Exports shows from module import * and
the role __all__ plays in it. Without
__all__, the star import falls back to a naming
convention. Copy exporting.py without the
line, add a name of your own, and print what dir()
reports after the import.
# The shape of exporting_no_all.py
def public():
...
def helper():
...
def _internal():
...
def undeclared():
...# exporting_no_all.py
def public():
return "public"
def helper():
return "helper"
def _internal():
return "internal"
def undeclared():
return "undeclared"# exercise_6.py
from exporting_no_all import * # noqa: F403
print(sorted(n for n in dir() if not n.startswith("__")))
#: ['helper', 'public', 'undeclared']Without __all__, the star import falls back to
the underscore convention: every top-level name that does not
start with an underscore arrives. undeclared
therefore joins public and helper, and
_internal stays out. Restoring the
__all__ line shrinks the surface back to
public and helper.
__all__ and the underscore convention compose in
one direction only: __all__ can export an
underscored name, but without __all__ an underscore
is the only way to keep a name out of a star import.
from import shares the object, not the nameGive a module a top-level list,
plugins = [], and bring the list into a script withfrom ... import plugins. Append to the list through the module’s name, then print the script’splugins. Rebind the module’s name to a new list, append to that one, and print both. Explain why the first change reaches the script’s name and the second does not.
In Importing
Names with from and as, from_snapshot.py shows an
imported name keeping the value it had at import time. A
from import binds your name to the object, not to
the module’s name. Compare mutating the list with rebinding the
module’s attribute, and use is to check whether the
two names still share one object.
# plugin_list.py
plugins = []# exercise_7.py
import plugin_list
from plugin_list import plugins
plugin_list.plugins.append("spell check")
print(plugins)
#: ['spell check']
print(plugins is plugin_list.plugins)
#: True
plugin_list.plugins = []
plugin_list.plugins.append("word count")
print(plugins, plugin_list.plugins)
#: ['spell check'] ['word count']
print(plugins is plugin_list.plugins)
#: FalseShare one list between two names.
from plugin_list import plugins binds the script’s
plugins to the list the module’s name references,
so at first the two names share one object. Appending changes
that object, and both names show the new item.
Replace the module’s list. The assignment
plugin_list.plugins = [] rebinds the module’s name
to a second list and leaves the script’s name on the first, so
the second append() reaches a list to which the
script’s plugins does not refer.
exercise_7.py is from_snapshot.py with a
mutable value: the from import takes no copy, and
it does not follow the module’s name when that name moves. When
a module’s name can move to a new list or dict, import the
module and read plugin_list.plugins each time.