codekingpro/portable-devtools
114k
1'''"Executable documentation" for the pickle module.2 3Extensive comments about the pickle protocols and pickle-machine opcodes4can be found here. Some functions meant for external use:5 6genops(pickle)7 Generate all the opcodes in a pickle, as (opcode, arg, position) triples.8 9dis(pickle, out=None, memo=None, indentlevel=4)10 Print a symbolic disassembly of a pickle.11'''12 13import codecs14import io15import pickle16import re17import sys18 19__all__ = ['dis', 'genops', 'optimize']20 21bytes_types = pickle.bytes_types22 23# Other ideas:24#25# - A pickle verifier: read a pickle and check it exhaustively for26# well-formedness. dis() does a lot of this already.27#28# - A protocol identifier: examine a pickle and return its protocol number29# (== the highest .proto attr value among all the opcodes in the pickle).30# dis() already prints this info at the end.31#32# - A pickle optimizer: for example, tuple-building code is sometimes more33# elaborate than necessary, catering for the possibility that the tuple34# is recursive. Or lots of times a PUT is generated that's never accessed35# by a later GET.36 37 38# "A pickle" is a program for a virtual pickle machine (PM, but more accurately39# called an unpickling machine). It's a sequence of opcodes, interpreted by the40# PM, building an arbitrarily complex Python object.41#42# For the most part, the PM is very simple: there are no looping, testing, or43# conditional instructions, no arithmetic and no function calls. Opcodes are44# executed once each, from first to last, until a STOP opcode is reached.45#46# The PM has two data areas, "the stack" and "the memo".47#48# Many opcodes push Python objects onto the stack; e.g., INT pushes a Python49# integer object on the stack, whose value is gotten from a decimal string50# literal immediately following the INT opcode in the pickle bytestream. Other51# opcodes take Python objects off the stack. The result of unpickling is52# whatever object is left on the stack when the final STOP opcode is executed.53#54# The memo is simply an array of objects, or it can be implemented as a dict55# mapping little integers to objects. The memo serves as the PM's "long term56# memory", and the little integers indexing the memo are akin to variable57# names. Some opcodes pop a stack object into the memo at a given index,58# and others push a memo object at a given index onto the stack again.59#60# At heart, that's all the PM has. Subtleties arise for these reasons:61#62# + Object identity. Objects can be arbitrarily complex, and subobjects63# may be shared (for example, the list [a, a] refers to the same object a64# twice). It can be vital that unpickling recreate an isomorphic object65# graph, faithfully reproducing sharing.66#67# + Recursive objects. For example, after "L = []; L.append(L)", L is a68# list, and L[0] is the same list. This is related to the object identity69# point, and some sequences of pickle opcodes are subtle in order to70# get the right result in all cases.71#72# + Things pickle doesn't know everything about. Examples of things pickle73# does know everything about are Python's builtin scalar and container74# types, like ints and tuples. They generally have opcodes dedicated to75# them. For things like module references and instances of user-defined76# classes, pickle's knowledge is limited. Historically, many enhancements77# have been made to the pickle protocol in order to do a better (faster,78# and/or more compact) job on those.79#80# + Backward compatibility and micro-optimization. As explained below,81# pickle opcodes never go away, not even when better ways to do a thing82# get invented. The repertoire of the PM just keeps growing over time.83# For example, protocol 0 had two opcodes for building Python integers (INT84# and LONG), protocol 1 added three more for more-efficient pickling of short85# integers, and protocol 2 added two more for more-efficient pickling of86# long integers (before protocol 2, the only ways to pickle a Python long87# took time quadratic in the number of digits, for both pickling and88# unpickling). "Opcode bloat" isn't so much a subtlety as a source of89# wearying complication.90#91#92# Pickle protocols:93#94# For compatibility, the meaning of a pickle opcode never changes. Instead new95# pickle opcodes get added, and each version's unpickler can handle all the96# pickle opcodes in all protocol versions to date. So old pickles continue to97# be readable forever. The pickler can generally be told to restrict itself to98# the subset of opcodes available under previous protocol versions too, so that99# users can create pickles under the current version readable by older100# versions. However, a pickle does not contain its version number embedded101# within it. If an older unpickler tries to read a pickle using a later102# protocol, the result is most likely an exception due to seeing an unknown (in103# the older unpickler) opcode.104#105# The original pickle used what's now called "protocol 0", and what was called106# "text mode" before Python 2.3. The entire pickle bytestream is made up of107# printable 7-bit ASCII characters, plus the newline character, in protocol 0.108# That's why it was called text mode. Protocol 0 is small and elegant, but109# sometimes painfully inefficient.110#111# The second major set of additions is now called "protocol 1", and was called112# "binary mode" before Python 2.3. This added many opcodes with arguments113# consisting of arbitrary bytes, including NUL bytes and unprintable "high bit"114# bytes. Binary mode pickles can be substantially smaller than equivalent115# text mode pickles, and sometimes faster too; e.g., BININT represents a 4-byte116# int as 4 bytes following the opcode, which is cheaper to unpickle than the117# (perhaps) 11-character decimal string attached to INT. Protocol 1 also added118# a number of opcodes that operate on many stack elements at once (like APPENDS119# and SETITEMS), and "shortcut" opcodes (like EMPTY_DICT and EMPTY_TUPLE).120#121# The third major set of additions came in Python 2.3, and is called "protocol122# 2". This added:123#124# - A better way to pickle instances of new-style classes (NEWOBJ).125#126# - A way for a pickle to identify its protocol (PROTO).127#128# - Time- and space- efficient pickling of long ints (LONG{1,4}).129#130# - Shortcuts for small tuples (TUPLE{1,2,3}}.131#132# - Dedicated opcodes for bools (NEWTRUE, NEWFALSE).133#134# - The "extension registry", a vector of popular objects that can be pushed135# efficiently by index (EXT{1,2,4}). This is akin to the memo and GET, but136# the registry contents are predefined (there's nothing akin to the memo's137# PUT).138#139# Another independent change with Python 2.3 is the abandonment of any140# pretense that it might be safe to load pickles received from untrusted141# parties -- no sufficient security analysis has been done to guarantee142# this and there isn't a use case that warrants the expense of such an143# analysis.144#145# To this end, all tests for __safe_for_unpickling__ or for146# copyreg.safe_constructors are removed from the unpickling code.147# References to these variables in the descriptions below are to be seen148# as describing unpickling in Python 2.2 and before.149 150 151# Meta-rule: Descriptions are stored in instances of descriptor objects,152# with plain constructors. No meta-language is defined from which153# descriptors could be constructed. If you want, e.g., XML, write a little154# program to generate XML from the objects.155 156##############################################################################157# Some pickle opcodes have an argument, following the opcode in the158# bytestream. An argument is of a specific type, described by an instance159# of ArgumentDescriptor. These are not to be confused with arguments taken160# off the stack -- ArgumentDescriptor applies only to arguments embedded in161# the opcode stream, immediately following an opcode.162 163# Represents the number of bytes consumed by an argument delimited by the164# next newline character.165UP_TO_NEWLINE = -1166 167# Represents the number of bytes consumed by a two-argument opcode where168# the first argument gives the number of bytes in the second argument.169TAKEN_FROM_ARGUMENT1 = -2 # num bytes is 1-byte unsigned int170TAKEN_FROM_ARGUMENT4 = -3 # num bytes is 4-byte signed little-endian int171TAKEN_FROM_ARGUMENT4U = -4 # num bytes is 4-byte unsigned little-endian int172TAKEN_FROM_ARGUMENT8U = -5 # num bytes is 8-byte unsigned little-endian int173 174class ArgumentDescriptor(object):175 __slots__ = (176 # name of descriptor record, also a module global name; a string177 'name',178 179 # length of argument, in bytes; an int; UP_TO_NEWLINE and180 # TAKEN_FROM_ARGUMENT{1,4,8} are negative values for variable-length181 # cases182 'n',183 184 # a function taking a file-like object, reading this kind of argument185 # from the object at the current position, advancing the current186 # position by n bytes, and returning the value of the argument187 'reader',188 189 # human-readable docs for this arg descriptor; a string190 'doc',191 )192 193 def __init__(self, name, n, reader, doc):194 assert isinstance(name, str)195 self.name = name196 197 assert isinstance(n, int) and (n >= 0 or198 n in (UP_TO_NEWLINE,199 TAKEN_FROM_ARGUMENT1,200 TAKEN_FROM_ARGUMENT4,201 TAKEN_FROM_ARGUMENT4U,202 TAKEN_FROM_ARGUMENT8U))203 self.n = n204 205 self.reader = reader206 207 assert isinstance(doc, str)208 self.doc = doc209 210from struct import unpack as _unpack211 212def read_uint1(f):213 r"""214 >>> import io215 >>> read_uint1(io.BytesIO(b'\xff'))216 255217 """218 219 data = f.read(1)220 if data:221 return data[0]222 raise ValueError("not enough data in stream to read uint1")223 224uint1 = ArgumentDescriptor(225 name='uint1',226 n=1,227 reader=read_uint1,228 doc="One-byte unsigned integer.")229 230 231def read_uint2(f):232 r"""233 >>> import io234 >>> read_uint2(io.BytesIO(b'\xff\x00'))235 255236 >>> read_uint2(io.BytesIO(b'\xff\xff'))237 65535238 """239 240 data = f.read(2)241 if len(data) == 2:242 return _unpack("<H", data)[0]243 raise ValueError("not enough data in stream to read uint2")244 245uint2 = ArgumentDescriptor(246 name='uint2',247 n=2,248 reader=read_uint2,249 doc="Two-byte unsigned integer, little-endian.")250 251 252def read_int4(f):253 r"""254 >>> import io255 >>> read_int4(io.BytesIO(b'\xff\x00\x00\x00'))256 255257 >>> read_int4(io.BytesIO(b'\x00\x00\x00\x80')) == -(2**31)258 True259 """260 261 data = f.read(4)262 if len(data) == 4:263 return _unpack("<i", data)[0]264 raise ValueError("not enough data in stream to read int4")265 266int4 = ArgumentDescriptor(267 name='int4',268 n=4,269 reader=read_int4,270 doc="Four-byte signed integer, little-endian, 2's complement.")271 272 273def read_uint4(f):274 r"""275 >>> import io276 >>> read_uint4(io.BytesIO(b'\xff\x00\x00\x00'))277 255278 >>> read_uint4(io.BytesIO(b'\x00\x00\x00\x80')) == 2**31279 True280 """281 282 data = f.read(4)283 if len(data) == 4:284 return _unpack("<I", data)[0]285 raise ValueError("not enough data in stream to read uint4")286 287uint4 = ArgumentDescriptor(288 name='uint4',289 n=4,290 reader=read_uint4,291 doc="Four-byte unsigned integer, little-endian.")292 293 294def read_uint8(f):295 r"""296 >>> import io297 >>> read_uint8(io.BytesIO(b'\xff\x00\x00\x00\x00\x00\x00\x00'))298 255299 >>> read_uint8(io.BytesIO(b'\xff' * 8)) == 2**64-1300 True301 """302 303 data = f.read(8)304 if len(data) == 8:305 return _unpack("<Q", data)[0]306 raise ValueError("not enough data in stream to read uint8")307 308uint8 = ArgumentDescriptor(309 name='uint8',310 n=8,311 reader=read_uint8,312 doc="Eight-byte unsigned integer, little-endian.")313 314 315def read_stringnl(f, decode=True, stripquotes=True, *, encoding='latin-1'):316 r"""317 >>> import io318 >>> read_stringnl(io.BytesIO(b"'abcd'\nefg\n"))319 'abcd'320 321 >>> read_stringnl(io.BytesIO(b"\n"))322 Traceback (most recent call last):323 ...324 ValueError: no string quotes around b''325 326 >>> read_stringnl(io.BytesIO(b"\n"), stripquotes=False)327 ''328 329 >>> read_stringnl(io.BytesIO(b"''\n"))330 ''331 332 >>> read_stringnl(io.BytesIO(b'"abcd"'))333 Traceback (most recent call last):334 ...335 ValueError: no newline found when trying to read stringnl336 337 Embedded escapes are undone in the result.338 >>> read_stringnl(io.BytesIO(br"'a\n\\b\x00c\td'" + b"\n'e'"))339 'a\n\\b\x00c\td'340 """341 342 data = f.readline()343 if not data.endswith(b'\n'):344 raise ValueError("no newline found when trying to read stringnl")345 data = data[:-1] # lose the newline346 347 if stripquotes:348 for q in (b'"', b"'"):349 if data.startswith(q):350 if not data.endswith(q):351 raise ValueError("string quote %r not found at both "352 "ends of %r" % (q, data))353 data = data[1:-1]354 break355 else:356 raise ValueError("no string quotes around %r" % data)357 358 if decode:359 data = codecs.escape_decode(data)[0].decode(encoding)360 return data361 362stringnl = ArgumentDescriptor(363 name='stringnl',364 n=UP_TO_NEWLINE,365 reader=read_stringnl,366 doc="""A newline-terminated string.367 368 This is a repr-style string, with embedded escapes, and369 bracketing quotes.370 """)371 372def read_stringnl_noescape(f):373 return read_stringnl(f, stripquotes=False, encoding='utf-8')374 375stringnl_noescape = ArgumentDescriptor(376 name='stringnl_noescape',377 n=UP_TO_NEWLINE,378 reader=read_stringnl_noescape,379 doc="""A newline-terminated string.380 381 This is a str-style string, without embedded escapes,382 or bracketing quotes. It should consist solely of383 printable ASCII characters.384 """)385 386def read_stringnl_noescape_pair(f):387 r"""388 >>> import io389 >>> read_stringnl_noescape_pair(io.BytesIO(b"Queue\nEmpty\njunk"))390 'Queue Empty'391 """392 393 return "%s %s" % (read_stringnl_noescape(f), read_stringnl_noescape(f))394 395stringnl_noescape_pair = ArgumentDescriptor(396 name='stringnl_noescape_pair',397 n=UP_TO_NEWLINE,398 reader=read_stringnl_noescape_pair,399 doc="""A pair of newline-terminated strings.400 401 These are str-style strings, without embedded402 escapes, or bracketing quotes. They should403 consist solely of printable ASCII characters.404 The pair is returned as a single string, with405 a single blank separating the two strings.406 """)407 408 409def read_string1(f):410 r"""411 >>> import io412 >>> read_string1(io.BytesIO(b"\x00"))413 ''414 >>> read_string1(io.BytesIO(b"\x03abcdef"))415 'abc'416 """417 418 n = read_uint1(f)419 assert n >= 0420 data = f.read(n)421 if len(data) == n:422 return data.decode("latin-1")423 raise ValueError("expected %d bytes in a string1, but only %d remain" %424 (n, len(data)))425 426string1 = ArgumentDescriptor(427 name="string1",428 n=TAKEN_FROM_ARGUMENT1,429 reader=read_string1,430 doc="""A counted string.431 432 The first argument is a 1-byte unsigned int giving the number433 of bytes in the string, and the second argument is that many434 bytes.435 """)436 437 438def read_string4(f):439 r"""440 >>> import io441 >>> read_string4(io.BytesIO(b"\x00\x00\x00\x00abc"))442 ''443 >>> read_string4(io.BytesIO(b"\x03\x00\x00\x00abcdef"))444 'abc'445 >>> read_string4(io.BytesIO(b"\x00\x00\x00\x03abcdef"))446 Traceback (most recent call last):447 ...448 ValueError: expected 50331648 bytes in a string4, but only 6 remain449 """450 451 n = read_int4(f)452 if n < 0:453 raise ValueError("string4 byte count < 0: %d" % n)454 data = f.read(n)455 if len(data) == n:456 return data.decode("latin-1")457 raise ValueError("expected %d bytes in a string4, but only %d remain" %458 (n, len(data)))459 460string4 = ArgumentDescriptor(461 name="string4",462 n=TAKEN_FROM_ARGUMENT4,463 reader=read_string4,464 doc="""A counted string.465 466 The first argument is a 4-byte little-endian signed int giving467 the number of bytes in the string, and the second argument is468 that many bytes.469 """)470 471 472def read_bytes1(f):473 r"""474 >>> import io475 >>> read_bytes1(io.BytesIO(b"\x00"))476 b''477 >>> read_bytes1(io.BytesIO(b"\x03abcdef"))478 b'abc'479 """480 481 n = read_uint1(f)482 assert n >= 0483 data = f.read(n)484 if len(data) == n:485 return data486 raise ValueError("expected %d bytes in a bytes1, but only %d remain" %487 (n, len(data)))488 489bytes1 = ArgumentDescriptor(490 name="bytes1",491 n=TAKEN_FROM_ARGUMENT1,492 reader=read_bytes1,493 doc="""A counted bytes string.494 495 The first argument is a 1-byte unsigned int giving the number496 of bytes, and the second argument is that many bytes.497 """)498 499 500def read_bytes4(f):501 r"""502 >>> import io503 >>> read_bytes4(io.BytesIO(b"\x00\x00\x00\x00abc"))504 b''505 >>> read_bytes4(io.BytesIO(b"\x03\x00\x00\x00abcdef"))506 b'abc'507 >>> read_bytes4(io.BytesIO(b"\x00\x00\x00\x03abcdef"))508 Traceback (most recent call last):509 ...510 ValueError: expected 50331648 bytes in a bytes4, but only 6 remain511 """512 513 n = read_uint4(f)514 assert n >= 0515 if n > sys.maxsize:516 raise ValueError("bytes4 byte count > sys.maxsize: %d" % n)517 data = f.read(n)518 if len(data) == n:519 return data520 raise ValueError("expected %d bytes in a bytes4, but only %d remain" %521 (n, len(data)))522 523bytes4 = ArgumentDescriptor(524 name="bytes4",525 n=TAKEN_FROM_ARGUMENT4U,526 reader=read_bytes4,527 doc="""A counted bytes string.528 529 The first argument is a 4-byte little-endian unsigned int giving530 the number of bytes, and the second argument is that many bytes.531 """)532 533 534def read_bytes8(f):535 r"""536 >>> import io, struct, sys537 >>> read_bytes8(io.BytesIO(b"\x00\x00\x00\x00\x00\x00\x00\x00abc"))538 b''539 >>> read_bytes8(io.BytesIO(b"\x03\x00\x00\x00\x00\x00\x00\x00abcdef"))540 b'abc'541 >>> bigsize8 = struct.pack("<Q", sys.maxsize//3)542 >>> read_bytes8(io.BytesIO(bigsize8 + b"abcdef")) #doctest: +ELLIPSIS543 Traceback (most recent call last):544 ...545 ValueError: expected ... bytes in a bytes8, but only 6 remain546 """547 548 n = read_uint8(f)549 assert n >= 0550 if n > sys.maxsize:551 raise ValueError("bytes8 byte count > sys.maxsize: %d" % n)552 data = f.read(n)553 if len(data) == n:554 return data555 raise ValueError("expected %d bytes in a bytes8, but only %d remain" %556 (n, len(data)))557 558bytes8 = ArgumentDescriptor(559 name="bytes8",560 n=TAKEN_FROM_ARGUMENT8U,561 reader=read_bytes8,562 doc="""A counted bytes string.563 564 The first argument is an 8-byte little-endian unsigned int giving565 the number of bytes, and the second argument is that many bytes.566 """)567 568 569def read_bytearray8(f):570 r"""571 >>> import io, struct, sys572 >>> read_bytearray8(io.BytesIO(b"\x00\x00\x00\x00\x00\x00\x00\x00abc"))573 bytearray(b'')574 >>> read_bytearray8(io.BytesIO(b"\x03\x00\x00\x00\x00\x00\x00\x00abcdef"))575 bytearray(b'abc')576 >>> bigsize8 = struct.pack("<Q", sys.maxsize//3)577 >>> read_bytearray8(io.BytesIO(bigsize8 + b"abcdef")) #doctest: +ELLIPSIS578 Traceback (most recent call last):579 ...580 ValueError: expected ... bytes in a bytearray8, but only 6 remain581 """582 583 n = read_uint8(f)584 assert n >= 0585 if n > sys.maxsize:586 raise ValueError("bytearray8 byte count > sys.maxsize: %d" % n)587 data = f.read(n)588 if len(data) == n:589 return bytearray(data)590 raise ValueError("expected %d bytes in a bytearray8, but only %d remain" %591 (n, len(data)))592 593bytearray8 = ArgumentDescriptor(594 name="bytearray8",595 n=TAKEN_FROM_ARGUMENT8U,596 reader=read_bytearray8,597 doc="""A counted bytearray.598 599 The first argument is an 8-byte little-endian unsigned int giving600 the number of bytes, and the second argument is that many bytes.601 """)602 603def read_unicodestringnl(f):604 r"""605 >>> import io606 >>> read_unicodestringnl(io.BytesIO(b"abc\\uabcd\njunk")) == 'abc\uabcd'607 True608 """609 610 data = f.readline()611 if not data.endswith(b'\n'):612 raise ValueError("no newline found when trying to read "613 "unicodestringnl")614 data = data[:-1] # lose the newline615 return str(data, 'raw-unicode-escape')616 617unicodestringnl = ArgumentDescriptor(618 name='unicodestringnl',619 n=UP_TO_NEWLINE,620 reader=read_unicodestringnl,621 doc="""A newline-terminated Unicode string.622 623 This is raw-unicode-escape encoded, so consists of624 printable ASCII characters, and may contain embedded625 escape sequences.626 """)627 628 629def read_unicodestring1(f):630 r"""631 >>> import io632 >>> s = 'abcd\uabcd'633 >>> enc = s.encode('utf-8')634 >>> enc635 b'abcd\xea\xaf\x8d'636 >>> n = bytes([len(enc)]) # little-endian 1-byte length637 >>> t = read_unicodestring1(io.BytesIO(n + enc + b'junk'))638 >>> s == t639 True640 641 >>> read_unicodestring1(io.BytesIO(n + enc[:-1]))642 Traceback (most recent call last):643 ...644 ValueError: expected 7 bytes in a unicodestring1, but only 6 remain645 """646 647 n = read_uint1(f)648 assert n >= 0649 data = f.read(n)650 if len(data) == n:651 return str(data, 'utf-8', 'surrogatepass')652 raise ValueError("expected %d bytes in a unicodestring1, but only %d "653 "remain" % (n, len(data)))654 655unicodestring1 = ArgumentDescriptor(656 name="unicodestring1",657 n=TAKEN_FROM_ARGUMENT1,658 reader=read_unicodestring1,659 doc="""A counted Unicode string.660 661 The first argument is a 1-byte little-endian signed int662 giving the number of bytes in the string, and the second663 argument-- the UTF-8 encoding of the Unicode string --664 contains that many bytes.665 """)666 667 668def read_unicodestring4(f):669 r"""670 >>> import io671 >>> s = 'abcd\uabcd'672 >>> enc = s.encode('utf-8')673 >>> enc674 b'abcd\xea\xaf\x8d'675 >>> n = bytes([len(enc), 0, 0, 0]) # little-endian 4-byte length676 >>> t = read_unicodestring4(io.BytesIO(n + enc + b'junk'))677 >>> s == t678 True679 680 >>> read_unicodestring4(io.BytesIO(n + enc[:-1]))681 Traceback (most recent call last):682 ...683 ValueError: expected 7 bytes in a unicodestring4, but only 6 remain684 """685 686 n = read_uint4(f)687 assert n >= 0688 if n > sys.maxsize:689 raise ValueError("unicodestring4 byte count > sys.maxsize: %d" % n)690 data = f.read(n)691 if len(data) == n:692 return str(data, 'utf-8', 'surrogatepass')693 raise ValueError("expected %d bytes in a unicodestring4, but only %d "694 "remain" % (n, len(data)))695 696unicodestring4 = ArgumentDescriptor(697 name="unicodestring4",698 n=TAKEN_FROM_ARGUMENT4U,699 reader=read_unicodestring4,700 doc="""A counted Unicode string.701 702 The first argument is a 4-byte little-endian signed int703 giving the number of bytes in the string, and the second704 argument-- the UTF-8 encoding of the Unicode string --705 contains that many bytes.706 """)707 708 709def read_unicodestring8(f):710 r"""711 >>> import io712 >>> s = 'abcd\uabcd'713 >>> enc = s.encode('utf-8')714 >>> enc715 b'abcd\xea\xaf\x8d'716 >>> n = bytes([len(enc)]) + b'\0' * 7 # little-endian 8-byte length717 >>> t = read_unicodestring8(io.BytesIO(n + enc + b'junk'))718 >>> s == t719 True720 721 >>> read_unicodestring8(io.BytesIO(n + enc[:-1]))722 Traceback (most recent call last):723 ...724 ValueError: expected 7 bytes in a unicodestring8, but only 6 remain725 """726 727 n = read_uint8(f)728 assert n >= 0729 if n > sys.maxsize:730 raise ValueError("unicodestring8 byte count > sys.maxsize: %d" % n)731 data = f.read(n)732 if len(data) == n:733 return str(data, 'utf-8', 'surrogatepass')734 raise ValueError("expected %d bytes in a unicodestring8, but only %d "735 "remain" % (n, len(data)))736 737unicodestring8 = ArgumentDescriptor(738 name="unicodestring8",739 n=TAKEN_FROM_ARGUMENT8U,740 reader=read_unicodestring8,741 doc="""A counted Unicode string.742 743 The first argument is an 8-byte little-endian signed int744 giving the number of bytes in the string, and the second745 argument-- the UTF-8 encoding of the Unicode string --746 contains that many bytes.747 """)748 749 750def read_decimalnl_short(f):751 r"""752 >>> import io753 >>> read_decimalnl_short(io.BytesIO(b"1234\n56"))754 1234755 756 >>> read_decimalnl_short(io.BytesIO(b"1234L\n56"))757 Traceback (most recent call last):758 ...759 ValueError: invalid literal for int() with base 10: b'1234L'760 """761 762 s = read_stringnl(f, decode=False, stripquotes=False)763 764 # There's a hack for True and False here.765 if s == b"00":766 return False767 elif s == b"01":768 return True769 770 return int(s)771 772def read_decimalnl_long(f):773 r"""774 >>> import io775 776 >>> read_decimalnl_long(io.BytesIO(b"1234L\n56"))777 1234778 779 >>> read_decimalnl_long(io.BytesIO(b"123456789012345678901234L\n6"))780 123456789012345678901234781 """782 783 s = read_stringnl(f, decode=False, stripquotes=False)784 if s[-1:] == b'L':785 s = s[:-1]786 return int(s)787 788 789decimalnl_short = ArgumentDescriptor(790 name='decimalnl_short',791 n=UP_TO_NEWLINE,792 reader=read_decimalnl_short,793 doc="""A newline-terminated decimal integer literal.794 795 This never has a trailing 'L', and the integer fit796 in a short Python int on the box where the pickle797 was written -- but there's no guarantee it will fit798 in a short Python int on the box where the pickle799 is read.800 """)801 802decimalnl_long = ArgumentDescriptor(803 name='decimalnl_long',804 n=UP_TO_NEWLINE,805 reader=read_decimalnl_long,806 doc="""A newline-terminated decimal integer literal.807 808 This has a trailing 'L', and can represent integers809 of any size.810 """)811 812 813def read_floatnl(f):814 r"""815 >>> import io816 >>> read_floatnl(io.BytesIO(b"-1.25\n6"))817 -1.25818 """819 s = read_stringnl(f, decode=False, stripquotes=False)820 return float(s)821 822floatnl = ArgumentDescriptor(823 name='floatnl',824 n=UP_TO_NEWLINE,825 reader=read_floatnl,826 doc="""A newline-terminated decimal floating literal.827 828 In general this requires 17 significant digits for roundtrip829 identity, and pickling then unpickling infinities, NaNs, and830 minus zero doesn't work across boxes, or on some boxes even831 on itself (e.g., Windows can't read the strings it produces832 for infinities or NaNs).833 """)834 835def read_float8(f):836 r"""837 >>> import io, struct838 >>> raw = struct.pack(">d", -1.25)839 >>> raw840 b'\xbf\xf4\x00\x00\x00\x00\x00\x00'841 >>> read_float8(io.BytesIO(raw + b"\n"))842 -1.25843 """844 845 data = f.read(8)846 if len(data) == 8:847 return _unpack(">d", data)[0]848 raise ValueError("not enough data in stream to read float8")849 850 851float8 = ArgumentDescriptor(852 name='float8',853 n=8,854 reader=read_float8,855 doc="""An 8-byte binary representation of a float, big-endian.856 857 The format is unique to Python, and shared with the struct858 module (format string '>d') "in theory" (the struct and pickle859 implementations don't share the code -- they should). It's860 strongly related to the IEEE-754 double format, and, in normal861 cases, is in fact identical to the big-endian 754 double format.862 On other boxes the dynamic range is limited to that of a 754863 double, and "add a half and chop" rounding is used to reduce864 the precision to 53 bits. However, even on a 754 box,865 infinities, NaNs, and minus zero may not be handled correctly866 (may not survive roundtrip pickling intact).867 """)868 869# Protocol 2 formats870 871from pickle import decode_long872 873def read_long1(f):874 r"""875 >>> import io876 >>> read_long1(io.BytesIO(b"\x00"))877 0878 >>> read_long1(io.BytesIO(b"\x02\xff\x00"))879 255880 >>> read_long1(io.BytesIO(b"\x02\xff\x7f"))881 32767882 >>> read_long1(io.BytesIO(b"\x02\x00\xff"))883 -256884 >>> read_long1(io.BytesIO(b"\x02\x00\x80"))885 -32768886 """887 888 n = read_uint1(f)889 data = f.read(n)890 if len(data) != n:891 raise ValueError("not enough data in stream to read long1")892 return decode_long(data)893 894long1 = ArgumentDescriptor(895 name="long1",896 n=TAKEN_FROM_ARGUMENT1,897 reader=read_long1,898 doc="""A binary long, little-endian, using 1-byte size.899 900 This first reads one byte as an unsigned size, then reads that901 many bytes and interprets them as a little-endian 2's-complement long.902 If the size is 0, that's taken as a shortcut for the long 0L.903 """)904 905def read_long4(f):906 r"""907 >>> import io908 >>> read_long4(io.BytesIO(b"\x02\x00\x00\x00\xff\x00"))909 255910 >>> read_long4(io.BytesIO(b"\x02\x00\x00\x00\xff\x7f"))911 32767912 >>> read_long4(io.BytesIO(b"\x02\x00\x00\x00\x00\xff"))913 -256914 >>> read_long4(io.BytesIO(b"\x02\x00\x00\x00\x00\x80"))915 -32768916 >>> read_long1(io.BytesIO(b"\x00\x00\x00\x00"))917 0918 """919 920 n = read_int4(f)921 if n < 0:922 raise ValueError("long4 byte count < 0: %d" % n)923 data = f.read(n)924 if len(data) != n:925 raise ValueError("not enough data in stream to read long4")926 return decode_long(data)927 928long4 = ArgumentDescriptor(929 name="long4",930 n=TAKEN_FROM_ARGUMENT4,931 reader=read_long4,932 doc="""A binary representation of a long, little-endian.933 934 This first reads four bytes as a signed size (but requires the935 size to be >= 0), then reads that many bytes and interprets them936 as a little-endian 2's-complement long. If the size is 0, that's taken937 as a shortcut for the int 0, although LONG1 should really be used938 then instead (and in any case where # of bytes < 256).939 """)940 941 942##############################################################################943# Object descriptors. The stack used by the pickle machine holds objects,944# and in the stack_before and stack_after attributes of OpcodeInfo945# descriptors we need names to describe the various types of objects that can946# appear on the stack.947 948class StackObject(object):949 __slots__ = (950 # name of descriptor record, for info only951 'name',952 953 # type of object, or tuple of type objects (meaning the object can954 # be of any type in the tuple)955 'obtype',956 957 # human-readable docs for this kind of stack object; a string958 'doc',959 )960 961 def __init__(self, name, obtype, doc):962 assert isinstance(name, str)963 self.name = name964 965 assert isinstance(obtype, type) or isinstance(obtype, tuple)966 if isinstance(obtype, tuple):967 for contained in obtype:968 assert isinstance(contained, type)969 self.obtype = obtype970 971 assert isinstance(doc, str)972 self.doc = doc973 974 def __repr__(self):975 return self.name976 977 978pyint = pylong = StackObject(979 name='int',980 obtype=int,981 doc="A Python integer object.")982 983pyinteger_or_bool = StackObject(984 name='int_or_bool',985 obtype=(int, bool),986 doc="A Python integer or boolean object.")987 988pybool = StackObject(989 name='bool',990 obtype=bool,991 doc="A Python boolean object.")992 993pyfloat = StackObject(994 name='float',995 obtype=float,996 doc="A Python float object.")997 998pybytes_or_str = pystring = StackObject(999 name='bytes_or_str',1000 obtype=(bytes, str),1001 doc="A Python bytes or (Unicode) string object.")1002 1003pybytes = StackObject(1004 name='bytes',1005 obtype=bytes,1006 doc="A Python bytes object.")1007 1008pybytearray = StackObject(1009 name='bytearray',1010 obtype=bytearray,1011 doc="A Python bytearray object.")1012 1013pyunicode = StackObject(1014 name='str',1015 obtype=str,1016 doc="A Python (Unicode) string object.")1017 1018pynone = StackObject(1019 name="None",1020 obtype=type(None),1021 doc="The Python None object.")1022 1023pytuple = StackObject(1024 name="tuple",1025 obtype=tuple,1026 doc="A Python tuple object.")1027 1028pylist = StackObject(1029 name="list",1030 obtype=list,1031 doc="A Python list object.")1032 1033pydict = StackObject(1034 name="dict",1035 obtype=dict,1036 doc="A Python dict object.")1037 1038pyset = StackObject(1039 name="set",1040 obtype=set,1041 doc="A Python set object.")1042 1043pyfrozenset = StackObject(1044 name="frozenset",1045 obtype=set,1046 doc="A Python frozenset object.")1047 1048pybuffer = StackObject(1049 name='buffer',1050 obtype=object,1051 doc="A Python buffer-like object.")1052 1053anyobject = StackObject(1054 name='any',1055 obtype=object,1056 doc="Any kind of object whatsoever.")1057 1058markobject = StackObject(1059 name="mark",1060 obtype=StackObject,1061 doc="""'The mark' is a unique object.1062 1063Opcodes that operate on a variable number of objects1064generally don't embed the count of objects in the opcode,1065or pull it off the stack. Instead the MARK opcode is used1066to push a special marker object on the stack, and then1067some other opcodes grab all the objects from the top of1068the stack down to (but not including) the topmost marker1069object.1070""")1071 1072stackslice = StackObject(1073 name="stackslice",1074 obtype=StackObject,1075 doc="""An object representing a contiguous slice of the stack.1076 1077This is used in conjunction with markobject, to represent all1078of the stack following the topmost markobject. For example,1079the POP_MARK opcode changes the stack from1080 1081 [..., markobject, stackslice]1082to1083 [...]1084 1085No matter how many object are on the stack after the topmost1086markobject, POP_MARK gets rid of all of them (including the1087topmost markobject too).1088""")1089 1090##############################################################################1091# Descriptors for pickle opcodes.1092 1093class OpcodeInfo(object):1094 1095 __slots__ = (1096 # symbolic name of opcode; a string1097 'name',1098 1099 # the code used in a bytestream to represent the opcode; a1100 # one-character string1101 'code',1102 1103 # If the opcode has an argument embedded in the byte string, an1104 # instance of ArgumentDescriptor specifying its type. Note that1105 # arg.reader(s) can be used to read and decode the argument from1106 # the bytestream s, and arg.doc documents the format of the raw1107 # argument bytes. If the opcode doesn't have an argument embedded1108 # in the bytestream, arg should be None.1109 'arg',1110 1111 # what the stack looks like before this opcode runs; a list1112 'stack_before',1113 1114 # what the stack looks like after this opcode runs; a list1115 'stack_after',1116 1117 # the protocol number in which this opcode was introduced; an int1118 'proto',1119 1120 # human-readable docs for this opcode; a string1121 'doc',1122 )1123 1124 def __init__(self, name, code, arg,1125 stack_before, stack_after, proto, doc):1126 assert isinstance(name, str)1127 self.name = name1128 1129 assert isinstance(code, str)1130 assert len(code) == 11131 self.code = code1132 1133 assert arg is None or isinstance(arg, ArgumentDescriptor)1134 self.arg = arg1135 1136 assert isinstance(stack_before, list)1137 for x in stack_before:1138 assert isinstance(x, StackObject)1139 self.stack_before = stack_before1140 1141 assert isinstance(stack_after, list)1142 for x in stack_after:1143 assert isinstance(x, StackObject)1144 self.stack_after = stack_after1145 1146 assert isinstance(proto, int) and 0 <= proto <= pickle.HIGHEST_PROTOCOL1147 self.proto = proto1148 1149 assert isinstance(doc, str)1150 self.doc = doc1151 1152I = OpcodeInfo1153opcodes = [1154 1155 # Ways to spell integers.1156 1157 I(name='INT',1158 code='I',1159 arg=decimalnl_short,1160 stack_before=[],1161 stack_after=[pyinteger_or_bool],1162 proto=0,1163 doc="""Push an integer or bool.1164 1165 The argument is a newline-terminated decimal literal string.1166 1167 The intent may have been that this always fit in a short Python int,1168 but INT can be generated in pickles written on a 64-bit box that1169 require a Python long on a 32-bit box. The difference between this1170 and LONG then is that INT skips a trailing 'L', and produces a short1171 int whenever possible.1172 1173 Another difference is due to that, when bool was introduced as a1174 distinct type in 2.3, builtin names True and False were also added to1175 2.2.2, mapping to ints 1 and 0. For compatibility in both directions,1176 True gets pickled as INT + "I01\\n", and False as INT + "I00\\n".1177 Leading zeroes are never produced for a genuine integer. The 2.31178 (and later) unpicklers special-case these and return bool instead;1179 earlier unpicklers ignore the leading "0" and return the int.1180 """),1181 1182 I(name='BININT',1183 code='J',1184 arg=int4,1185 stack_before=[],1186 stack_after=[pyint],1187 proto=1,1188 doc="""Push a four-byte signed integer.1189 1190 This handles the full range of Python (short) integers on a 32-bit1191 box, directly as binary bytes (1 for the opcode and 4 for the integer).1192 If the integer is non-negative and fits in 1 or 2 bytes, pickling via1193 BININT1 or BININT2 saves space.1194 """),1195 1196 I(name='BININT1',1197 code='K',1198 arg=uint1,1199 stack_before=[],1200 stack_after=[pyint],