parthtamu/rag-code-assistant
0
1<!DOCTYPE html>2 3<html lang="en" data-content_root="../">4 <head>5 <meta charset="utf-8" />6 <meta name="viewport" content="width=device-width, initial-scale=1.0" /><meta name="viewport" content="width=device-width, initial-scale=1" />7<meta property="og:title" content="re — Regular expression operations" />8<meta property="og:type" content="website" />9<meta property="og:url" content="https://docs.python.org/3/library/re.html" />10<meta property="og:site_name" content="Python documentation" />11<meta property="og:description" content="Source code: Lib/re/ This module provides regular expression matching operations similar to those found in Perl. Both patterns and strings to be searched can be Unicode strings ( str) as well as 8-..." />12<meta property="og:image:width" content="1146" />13<meta property="og:image:height" content="600" />14<meta property="og:image" content="https://docs.python.org/3.15/_images/social_previews/summary_library_re_49568c13.png" />15<meta property="og:image:alt" content="Source code: Lib/re/ This module provides regular expression matching operations similar to those found in Perl. Both patterns and strings to be searched can be Unicode strings ( str) as well as 8-..." />16<meta name="description" content="Source code: Lib/re/ This module provides regular expression matching operations similar to those found in Perl. Both patterns and strings to be searched can be Unicode strings ( str) as well as 8-..." />17<meta name="twitter:card" content="summary_large_image" />18<meta name="theme-color" content="#3776ab">19 20 <title>re — Regular expression operations — Python 3.15.0a6 documentation</title><meta name="viewport" content="width=device-width, initial-scale=1.0">21 22 <link rel="stylesheet" type="text/css" href="../_static/pygments.css?v=b86133f3" />23 <link rel="stylesheet" type="text/css" href="../_static/classic.css?v=234b1a7c" />24 <link rel="stylesheet" type="text/css" href="../_static/pydoctheme.css?v=89a2f22a" />25 <link rel="stylesheet" type="text/css" href="../_static/profiling-sampling-visualization.css?v=0c2600ae" />26 <link id="pygments_dark_css" media="(prefers-color-scheme: dark)" rel="stylesheet" type="text/css" href="../_static/pygments_dark.css?v=5349f25f" />27 28 <script src="../_static/documentation_options.js?v=6b7c9ff5"></script>29 <script src="../_static/doctools.js?v=9bcbadda"></script>30 <script src="../_static/sphinx_highlight.js?v=dc90522c"></script>31 <script src="../_static/profiling-sampling-visualization.js?v=9811ed04"></script>32 33 <script src="../_static/sidebar.js"></script>34 35 <link rel="search" type="application/opensearchdescription+xml"36 title="Search within Python 3.15.0a6 documentation"37 href="../_static/opensearch.xml"/>38 <link rel="author" title="About these documents" href="../about.html" />39 <link rel="index" title="Index" href="../genindex.html" />40 <link rel="search" title="Search" href="../search.html" />41 <link rel="copyright" title="Copyright" href="../copyright.html" />42 <link rel="next" title="difflib — Helpers for computing deltas" href="difflib.html" />43 <link rel="prev" title="string.templatelib — Support for template string literals" href="string.templatelib.html" />44 45 46 <script defer file-types="bz2,epub,zip" data-domain="docs.python.org" src="https://analytics.python.org/js/script.file-downloads.outbound-links.js"></script>47 48 <link rel="canonical" href="https://docs.python.org/3/library/re.html">49 50 51 52 53 <style>54 @media only screen {55 table.full-width-table {56 width: 100%;57 }58 }59 </style>60<link rel="stylesheet" href="../_static/pydoctheme_dark.css" media="(prefers-color-scheme: dark)" id="pydoctheme_dark_css">61 <link rel="shortcut icon" type="image/png" href="../_static/py.svg">62 <script type="text/javascript" src="../_static/copybutton.js"></script>63 <script type="text/javascript" src="../_static/menu.js"></script>64 <script type="text/javascript" src="../_static/search-focus.js"></script>65 <script type="text/javascript" src="../_static/themetoggle.js"></script> 66 <script type="text/javascript" src="../_static/rtd_switcher.js"></script>67 <meta name="readthedocs-addons-api-version" content="1">68 69 </head>70<body>71<div class="mobile-nav">72 <input type="checkbox" id="menuToggler" class="toggler__input" aria-controls="navigation"73 aria-pressed="false" aria-expanded="false" role="button" aria-label="Menu">74 <nav class="nav-content" role="navigation">75 <label for="menuToggler" class="toggler__label">76 <span></span>77 </label>78 <span class="nav-items-wrapper">79 <a href="https://www.python.org/" class="nav-logo">80 <img src="../_static/py.svg" alt="Python logo">81 </a>82 <span class="version_switcher_placeholder"></span>83 <form role="search" class="search" action="../search.html" method="get">84 <svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" class="search-icon">85 <path fill-rule="nonzero" fill="currentColor" d="M15.5 14h-.79l-.28-.27a6.5 6.5 0 001.48-5.34c-.47-2.78-2.79-5-5.59-5.34a6.505 6.505 0 00-7.27 7.27c.34 2.8 2.56 5.12 5.34 5.59a6.5 6.5 0 005.34-1.48l.27.28v.79l4.25 4.25c.41.41 1.08.41 1.49 0 .41-.41.41-1.08 0-1.49L15.5 14zm-6 0C7.01 14 5 11.99 5 9.5S7.01 5 9.5 5 14 7.01 14 9.5 11.99 14 9.5 14z"></path>86 </svg>87 <input placeholder="Quick search" aria-label="Quick search" type="search" name="q">88 <input type="submit" value="Go">89 </form>90 </span>91 </nav>92 <div class="menu-wrapper">93 <nav class="menu" role="navigation" aria-label="main navigation">94 <div class="language_switcher_placeholder"></div>95 96<label class="theme-selector-label">97 Theme98 <select class="theme-selector" oninput="activateTheme(this.value)">99 <option value="auto" selected>Auto</option>100 <option value="light">Light</option>101 <option value="dark">Dark</option>102 </select>103</label>104 <div>105 <h3><a href="../contents.html">Table of Contents</a></h3>106 <ul>107<li><a class="reference internal" href="#"><code class="xref py py-mod docutils literal notranslate"><span class="pre">re</span></code> — Regular expression operations</a><ul>108<li><a class="reference internal" href="#regular-expression-syntax">Regular Expression Syntax</a></li>109<li><a class="reference internal" href="#module-contents">Module Contents</a><ul>110<li><a class="reference internal" href="#flags">Flags</a></li>111<li><a class="reference internal" href="#functions">Functions</a></li>112<li><a class="reference internal" href="#exceptions">Exceptions</a></li>113</ul>114</li>115<li><a class="reference internal" href="#regular-expression-objects">Regular Expression Objects</a></li>116<li><a class="reference internal" href="#match-objects">Match Objects</a></li>117<li><a class="reference internal" href="#regular-expression-examples">Regular Expression Examples</a><ul>118<li><a class="reference internal" href="#checking-for-a-pair">Checking for a Pair</a></li>119<li><a class="reference internal" href="#simulating-scanf">Simulating scanf()</a></li>120<li><a class="reference internal" href="#search-vs-prefixmatch">search() vs. prefixmatch()</a></li>121<li><a class="reference internal" href="#prefixmatch-vs-match">prefixmatch() vs. match()</a></li>122<li><a class="reference internal" href="#making-a-phonebook">Making a Phonebook</a></li>123<li><a class="reference internal" href="#text-munging">Text Munging</a></li>124<li><a class="reference internal" href="#finding-all-adverbs">Finding all Adverbs</a></li>125<li><a class="reference internal" href="#finding-all-adverbs-and-their-positions">Finding all Adverbs and their Positions</a></li>126<li><a class="reference internal" href="#raw-string-notation">Raw String Notation</a></li>127<li><a class="reference internal" href="#writing-a-tokenizer">Writing a Tokenizer</a></li>128</ul>129</li>130</ul>131</li>132</ul>133 134 </div>135 <div>136 <h4>Previous topic</h4>137 <p class="topless"><a href="string.templatelib.html"138 title="previous chapter"><code class="xref py py-mod docutils literal notranslate"><span class="pre">string.templatelib</span></code> — Support for template string literals</a></p>139 </div>140 <div>141 <h4>Next topic</h4>142 <p class="topless"><a href="difflib.html"143 title="next chapter"><code class="xref py py-mod docutils literal notranslate"><span class="pre">difflib</span></code> — Helpers for computing deltas</a></p>144 </div>145 <script>146 document.addEventListener('DOMContentLoaded', () => {147 const title = document.querySelector('meta[property="og:title"]').content;148 const elements = document.querySelectorAll('.improvepage');149 const pageurl = window.location.href.split('?')[0];150 elements.forEach(element => {151 const url = new URL(element.href.split('?')[0].replace("-nojs", ""));152 url.searchParams.set('pagetitle', title);153 url.searchParams.set('pageurl', pageurl);154 url.searchParams.set('pagesource', "library/re.rst");155 element.href = url.toString();156 });157 });158 </script>159 <div role="note" aria-label="source link">160 <h3>This page</h3>161 <ul class="this-page-menu">162 <li><a href="../bugs.html">Report a bug</a></li>163 <li><a class="improvepage" href="../improve-page-nojs.html">Improve this page</a></li>164 <li>165 <a href="https://github.com/python/cpython/blob/main/Doc/library/re.rst?plain=1"166 rel="nofollow">Show source167 </a>168 </li>169 170 </ul>171 </div>172 </nav>173 </div>174</div>175 176 177 <div class="related" role="navigation" aria-label="Related">178 <h3>Navigation</h3>179 <ul>180 <li class="right" style="margin-right: 10px">181 <a href="../genindex.html" title="General Index"182 accesskey="I">index</a></li>183 <li class="right" >184 <a href="../py-modindex.html" title="Python Module Index"185 >modules</a> |</li>186 <li class="right" >187 <a href="difflib.html" title="difflib — Helpers for computing deltas"188 accesskey="N">next</a> |</li>189 <li class="right" >190 <a href="string.templatelib.html" title="string.templatelib — Support for template string literals"191 accesskey="P">previous</a> |</li>192 193 <li><img src="../_static/py.svg" alt="Python logo" style="vertical-align: middle; margin-top: -1px"></li>194 <li><a href="https://www.python.org/">Python</a> »</li>195 <li class="switchers">196 <div class="language_switcher_placeholder"></div>197 <div class="version_switcher_placeholder"></div>198 </li>199 <li>200 201 </li>202 <li id="cpython-language-and-version">203 <a href="../index.html">3.15.0a6 Documentation</a> »204 </li>205 206 <li class="nav-item nav-item-1"><a href="index.html" >The Python Standard Library</a> »</li>207 <li class="nav-item nav-item-2"><a href="text.html" accesskey="U">Text Processing Services</a> »</li>208 <li class="nav-item nav-item-this"><a href=""><code class="xref py py-mod docutils literal notranslate"><span class="pre">re</span></code> — Regular expression operations</a></li>209 <li class="right">210 211 212 <div class="inline-search" role="search">213 <form class="inline-search" action="../search.html" method="get">214 <input placeholder="Quick search" aria-label="Quick search" type="search" name="q" id="search-box">215 <input type="submit" value="Go">216 </form>217 </div>218 |219 </li>220 <li class="right">221<label class="theme-selector-label">222 Theme223 <select class="theme-selector" oninput="activateTheme(this.value)">224 <option value="auto" selected>Auto</option>225 <option value="light">Light</option>226 <option value="dark">Dark</option>227 </select>228</label> |</li>229 230 </ul>231 </div> 232 233 <div class="document">234 <div class="documentwrapper">235 <div class="bodywrapper">236 <div class="body" role="main">237 238 <section id="module-re">239<span id="re-regular-expression-operations"></span><h1><code class="xref py py-mod docutils literal notranslate"><span class="pre">re</span></code> — Regular expression operations<a class="headerlink" href="#module-re" title="Link to this heading">¶</a></h1>240<p><strong>Source code:</strong> <a class="extlink-source reference external" href="https://github.com/python/cpython/tree/main/Lib/re/">Lib/re/</a></p>241<hr class="docutils" />242<p>This module provides regular expression matching operations similar to243those found in Perl.</p>244<p>Both patterns and strings to be searched can be Unicode strings (<a class="reference internal" href="stdtypes.html#str" title="str"><code class="xref py py-class docutils literal notranslate"><span class="pre">str</span></code></a>)245as well as 8-bit strings (<a class="reference internal" href="stdtypes.html#bytes" title="bytes"><code class="xref py py-class docutils literal notranslate"><span class="pre">bytes</span></code></a>).246However, Unicode strings and 8-bit strings cannot be mixed:247that is, you cannot match a Unicode string with a bytes pattern or248vice-versa; similarly, when asking for a substitution, the replacement249string must be of the same type as both the pattern and the search string.</p>250<p>Regular expressions use the backslash character (<code class="docutils literal notranslate"><span class="pre">'\'</span></code>) to indicate251special forms or to allow special characters to be used without invoking252their special meaning. This collides with Python’s usage of the same253character for the same purpose in string literals; for example, to match254a literal backslash, one might have to write <code class="docutils literal notranslate"><span class="pre">'\\\\'</span></code> as the pattern255string, because the regular expression must be <code class="docutils literal notranslate"><span class="pre">\\</span></code>, and each256backslash must be expressed as <code class="docutils literal notranslate"><span class="pre">\\</span></code> inside a regular Python string257literal. Also, please note that any invalid escape sequences in Python’s258usage of the backslash in string literals now generate a <a class="reference internal" href="exceptions.html#SyntaxWarning" title="SyntaxWarning"><code class="xref py py-exc docutils literal notranslate"><span class="pre">SyntaxWarning</span></code></a>259and in the future this will become a <a class="reference internal" href="exceptions.html#SyntaxError" title="SyntaxError"><code class="xref py py-exc docutils literal notranslate"><span class="pre">SyntaxError</span></code></a>. This behaviour260will happen even if it is a valid escape sequence for a regular expression.</p>261<p>The solution is to use Python’s raw string notation for regular expression262patterns; backslashes are not handled in any special way in a string literal263prefixed with <code class="docutils literal notranslate"><span class="pre">'r'</span></code>. So <code class="docutils literal notranslate"><span class="pre">r"\n"</span></code> is a two-character string containing264<code class="docutils literal notranslate"><span class="pre">'\'</span></code> and <code class="docutils literal notranslate"><span class="pre">'n'</span></code>, while <code class="docutils literal notranslate"><span class="pre">"\n"</span></code> is a one-character string containing a265newline. Usually patterns will be expressed in Python code using this raw266string notation.</p>267<p>It is important to note that most regular expression operations are available as268module-level functions and methods on269<a class="reference internal" href="#re-objects"><span class="std std-ref">compiled regular expressions</span></a>. The functions are shortcuts270that don’t require you to compile a regex object first, but miss some271fine-tuning parameters.</p>272<div class="admonition seealso">273<p class="admonition-title">See also</p>274<p>The third-party <a class="extlink-pypi reference external" href="https://pypi.org/project/regex/">regex</a> module,275which has an API compatible with the standard library <code class="xref py py-mod docutils literal notranslate"><span class="pre">re</span></code> module,276but offers additional functionality and a more thorough Unicode support.</p>277</div>278<section id="regular-expression-syntax">279<span id="re-syntax"></span><h2>Regular Expression Syntax<a class="headerlink" href="#regular-expression-syntax" title="Link to this heading">¶</a></h2>280<p>A regular expression (or RE) specifies a set of strings that matches it; the281functions in this module let you check if a particular string matches a given282regular expression (or if a given regular expression matches a particular283string, which comes down to the same thing).</p>284<p>Regular expressions can be concatenated to form new regular expressions; if <em>A</em>285and <em>B</em> are both regular expressions, then <em>AB</em> is also a regular expression.286In general, if a string <em>p</em> matches <em>A</em> and another string <em>q</em> matches <em>B</em>, the287string <em>pq</em> will match AB. This holds unless <em>A</em> or <em>B</em> contain low precedence288operations; boundary conditions between <em>A</em> and <em>B</em>; or have numbered group289references. Thus, complex expressions can easily be constructed from simpler290primitive expressions like the ones described here. For details of the theory291and implementation of regular expressions, consult the Friedl book <a class="reference internal" href="#frie09" id="id1"><span>[Frie09]</span></a>,292or almost any textbook about compiler construction.</p>293<p>A brief explanation of the format of regular expressions follows. For further294information and a gentler presentation, consult the <a class="reference internal" href="../howto/regex.html#regex-howto"><span class="std std-ref">Regular Expression HOWTO</span></a>.</p>295<p>Regular expressions can contain both special and ordinary characters. Most296ordinary characters, like <code class="docutils literal notranslate"><span class="pre">'A'</span></code>, <code class="docutils literal notranslate"><span class="pre">'a'</span></code>, or <code class="docutils literal notranslate"><span class="pre">'0'</span></code>, are the simplest regular297expressions; they simply match themselves. You can concatenate ordinary298characters, so <code class="docutils literal notranslate"><span class="pre">last</span></code> matches the string <code class="docutils literal notranslate"><span class="pre">'last'</span></code>. (In the rest of this299section, we’ll write RE’s in <code class="docutils literal notranslate"><span class="pre">this</span> <span class="pre">special</span> <span class="pre">style</span></code>, usually without quotes, and300strings to be matched <code class="docutils literal notranslate"><span class="pre">'in</span> <span class="pre">single</span> <span class="pre">quotes'</span></code>.)</p>301<p>Some characters, like <code class="docutils literal notranslate"><span class="pre">'|'</span></code> or <code class="docutils literal notranslate"><span class="pre">'('</span></code>, are special. Special302characters either stand for classes of ordinary characters, or affect303how the regular expressions around them are interpreted.</p>304<p>Repetition operators or quantifiers (<code class="docutils literal notranslate"><span class="pre">*</span></code>, <code class="docutils literal notranslate"><span class="pre">+</span></code>, <code class="docutils literal notranslate"><span class="pre">?</span></code>, <code class="docutils literal notranslate"><span class="pre">{m,n}</span></code>, etc) cannot be305directly nested. This avoids ambiguity with the non-greedy modifier suffix306<code class="docutils literal notranslate"><span class="pre">?</span></code>, and with other modifiers in other implementations. To apply a second307repetition to an inner repetition, parentheses may be used. For example,308the expression <code class="docutils literal notranslate"><span class="pre">(?:a{6})*</span></code> matches any multiple of six <code class="docutils literal notranslate"><span class="pre">'a'</span></code> characters.</p>309<p>The special characters are:</p>310<dl class="simple" id="index-0">311<dt><code class="docutils literal notranslate"><span class="pre">.</span></code></dt><dd><p>(Dot.) In the default mode, this matches any character except a newline. If312the <a class="reference internal" href="#re.DOTALL" title="re.DOTALL"><code class="xref py py-const docutils literal notranslate"><span class="pre">DOTALL</span></code></a> flag has been specified, this matches any character313including a newline. <code class="docutils literal notranslate"><span class="pre">(?s:.)</span></code> matches any character regardless of flags.</p>314</dd>315</dl>316<dl class="simple" id="index-1">317<dt><code class="docutils literal notranslate"><span class="pre">^</span></code></dt><dd><p>(Caret.) Matches the start of the string, and in <a class="reference internal" href="#re.MULTILINE" title="re.MULTILINE"><code class="xref py py-const docutils literal notranslate"><span class="pre">MULTILINE</span></code></a> mode also318matches immediately after each newline.</p>319</dd>320</dl>321<dl class="simple" id="index-2">322<dt><code class="docutils literal notranslate"><span class="pre">$</span></code></dt><dd><p>Matches the end of the string or just before the newline at the end of the323string, and in <a class="reference internal" href="#re.MULTILINE" title="re.MULTILINE"><code class="xref py py-const docutils literal notranslate"><span class="pre">MULTILINE</span></code></a> mode also matches before a newline. <code class="docutils literal notranslate"><span class="pre">foo</span></code>324matches both ‘foo’ and ‘foobar’, while the regular expression <code class="docutils literal notranslate"><span class="pre">foo$</span></code> matches325only ‘foo’. More interestingly, searching for <code class="docutils literal notranslate"><span class="pre">foo.$</span></code> in <code class="docutils literal notranslate"><span class="pre">'foo1\nfoo2\n'</span></code>326matches ‘foo2’ normally, but ‘foo1’ in <code class="xref py py-const docutils literal notranslate"><span class="pre">MULTILINE</span></code> mode; searching for327a single <code class="docutils literal notranslate"><span class="pre">$</span></code> in <code class="docutils literal notranslate"><span class="pre">'foo\n'</span></code> will find two (empty) matches: one just before328the newline, and one at the end of the string.</p>329</dd>330</dl>331<dl class="simple" id="index-3">332<dt><code class="docutils literal notranslate"><span class="pre">*</span></code></dt><dd><p>Causes the resulting RE to match 0 or more repetitions of the preceding RE, as333many repetitions as are possible. <code class="docutils literal notranslate"><span class="pre">ab*</span></code> will match ‘a’, ‘ab’, or ‘a’ followed334by any number of ‘b’s.</p>335</dd>336</dl>337<dl class="simple" id="index-4">338<dt><code class="docutils literal notranslate"><span class="pre">+</span></code></dt><dd><p>Causes the resulting RE to match 1 or more repetitions of the preceding RE.339<code class="docutils literal notranslate"><span class="pre">ab+</span></code> will match ‘a’ followed by any non-zero number of ‘b’s; it will not340match just ‘a’.</p>341</dd>342</dl>343<dl class="simple" id="index-5">344<dt><code class="docutils literal notranslate"><span class="pre">?</span></code></dt><dd><p>Causes the resulting RE to match 0 or 1 repetitions of the preceding RE.345<code class="docutils literal notranslate"><span class="pre">ab?</span></code> will match either ‘a’ or ‘ab’.</p>346</dd>347</dl>348<dl class="simple" id="index-6">349<dt><code class="docutils literal notranslate"><span class="pre">*?</span></code>, <code class="docutils literal notranslate"><span class="pre">+?</span></code>, <code class="docutils literal notranslate"><span class="pre">??</span></code></dt><dd><p>The <code class="docutils literal notranslate"><span class="pre">'*'</span></code>, <code class="docutils literal notranslate"><span class="pre">'+'</span></code>, and <code class="docutils literal notranslate"><span class="pre">'?'</span></code> quantifiers are all <em class="dfn">greedy</em>; they match350as much text as possible. Sometimes this behaviour isn’t desired; if the RE351<code class="docutils literal notranslate"><span class="pre"><.*></span></code> is matched against <code class="docutils literal notranslate"><span class="pre">'<a></span> <span class="pre">b</span> <span class="pre"><c>'</span></code>, it will match the entire352string, and not just <code class="docutils literal notranslate"><span class="pre">'<a>'</span></code>. Adding <code class="docutils literal notranslate"><span class="pre">?</span></code> after the quantifier makes it353perform the match in <em class="dfn">non-greedy</em> or <em class="dfn">minimal</em> fashion; as <em>few</em>354characters as possible will be matched. Using the RE <code class="docutils literal notranslate"><span class="pre"><.*?></span></code> will match355only <code class="docutils literal notranslate"><span class="pre">'<a>'</span></code>.</p>356</dd>357</dl>358<dl id="index-7">359<dt><code class="docutils literal notranslate"><span class="pre">*+</span></code>, <code class="docutils literal notranslate"><span class="pre">++</span></code>, <code class="docutils literal notranslate"><span class="pre">?+</span></code></dt><dd><p>Like the <code class="docutils literal notranslate"><span class="pre">'*'</span></code>, <code class="docutils literal notranslate"><span class="pre">'+'</span></code>, and <code class="docutils literal notranslate"><span class="pre">'?'</span></code> quantifiers, those where <code class="docutils literal notranslate"><span class="pre">'+'</span></code> is360appended also match as many times as possible.361However, unlike the true greedy quantifiers, these do not allow362back-tracking when the expression following it fails to match.363These are known as <em class="dfn">possessive</em> quantifiers.364For example, <code class="docutils literal notranslate"><span class="pre">a*a</span></code> will match <code class="docutils literal notranslate"><span class="pre">'aaaa'</span></code> because the <code class="docutils literal notranslate"><span class="pre">a*</span></code> will match365all 4 <code class="docutils literal notranslate"><span class="pre">'a'</span></code>s, but, when the final <code class="docutils literal notranslate"><span class="pre">'a'</span></code> is encountered, the366expression is backtracked so that in the end the <code class="docutils literal notranslate"><span class="pre">a*</span></code> ends up matching3673 <code class="docutils literal notranslate"><span class="pre">'a'</span></code>s total, and the fourth <code class="docutils literal notranslate"><span class="pre">'a'</span></code> is matched by the final <code class="docutils literal notranslate"><span class="pre">'a'</span></code>.368However, when <code class="docutils literal notranslate"><span class="pre">a*+a</span></code> is used to match <code class="docutils literal notranslate"><span class="pre">'aaaa'</span></code>, the <code class="docutils literal notranslate"><span class="pre">a*+</span></code> will369match all 4 <code class="docutils literal notranslate"><span class="pre">'a'</span></code>, but when the final <code class="docutils literal notranslate"><span class="pre">'a'</span></code> fails to find any more370characters to match, the expression cannot be backtracked and will thus371fail to match.372<code class="docutils literal notranslate"><span class="pre">x*+</span></code>, <code class="docutils literal notranslate"><span class="pre">x++</span></code> and <code class="docutils literal notranslate"><span class="pre">x?+</span></code> are equivalent to <code class="docutils literal notranslate"><span class="pre">(?>x*)</span></code>, <code class="docutils literal notranslate"><span class="pre">(?>x+)</span></code>373and <code class="docutils literal notranslate"><span class="pre">(?>x?)</span></code> correspondingly.</p>374<div class="versionadded">375<p><span class="versionmodified added">Added in version 3.11.</span></p>376</div>377</dd>378</dl>379<dl id="index-8">380<dt><code class="docutils literal notranslate"><span class="pre">{m}</span></code></dt><dd><p>Specifies that exactly <em>m</em> copies of the previous RE should be matched; fewer381matches cause the entire RE not to match. For example, <code class="docutils literal notranslate"><span class="pre">a{6}</span></code> will match382exactly six <code class="docutils literal notranslate"><span class="pre">'a'</span></code> characters, but not five.</p>383</dd>384<dt><code class="docutils literal notranslate"><span class="pre">{m,n}</span></code></dt><dd><p>Causes the resulting RE to match from <em>m</em> to <em>n</em> repetitions of the preceding385RE, attempting to match as many repetitions as possible. For example,386<code class="docutils literal notranslate"><span class="pre">a{3,5}</span></code> will match from 3 to 5 <code class="docutils literal notranslate"><span class="pre">'a'</span></code> characters. Omitting <em>m</em> specifies a387lower bound of zero, and omitting <em>n</em> specifies an infinite upper bound. As an388example, <code class="docutils literal notranslate"><span class="pre">a{4,}b</span></code> will match <code class="docutils literal notranslate"><span class="pre">'aaaab'</span></code> or a thousand <code class="docutils literal notranslate"><span class="pre">'a'</span></code> characters389followed by a <code class="docutils literal notranslate"><span class="pre">'b'</span></code>, but not <code class="docutils literal notranslate"><span class="pre">'aaab'</span></code>. The comma may not be omitted or the390modifier would be confused with the previously described form.</p>391</dd>392<dt><code class="docutils literal notranslate"><span class="pre">{m,n}?</span></code></dt><dd><p>Causes the resulting RE to match from <em>m</em> to <em>n</em> repetitions of the preceding393RE, attempting to match as <em>few</em> repetitions as possible. This is the394non-greedy version of the previous quantifier. For example, on the3956-character string <code class="docutils literal notranslate"><span class="pre">'aaaaaa'</span></code>, <code class="docutils literal notranslate"><span class="pre">a{3,5}</span></code> will match 5 <code class="docutils literal notranslate"><span class="pre">'a'</span></code> characters,396while <code class="docutils literal notranslate"><span class="pre">a{3,5}?</span></code> will only match 3 characters.</p>397</dd>398<dt><code class="docutils literal notranslate"><span class="pre">{m,n}+</span></code></dt><dd><p>Causes the resulting RE to match from <em>m</em> to <em>n</em> repetitions of the399preceding RE, attempting to match as many repetitions as possible400<em>without</em> establishing any backtracking points.401This is the possessive version of the quantifier above.402For example, on the 6-character string <code class="docutils literal notranslate"><span class="pre">'aaaaaa'</span></code>, <code class="docutils literal notranslate"><span class="pre">a{3,5}+aa</span></code>403attempt to match 5 <code class="docutils literal notranslate"><span class="pre">'a'</span></code> characters, then, requiring 2 more <code class="docutils literal notranslate"><span class="pre">'a'</span></code>s,404will need more characters than available and thus fail, while405<code class="docutils literal notranslate"><span class="pre">a{3,5}aa</span></code> will match with <code class="docutils literal notranslate"><span class="pre">a{3,5}</span></code> capturing 5, then 4 <code class="docutils literal notranslate"><span class="pre">'a'</span></code>s406by backtracking and then the final 2 <code class="docutils literal notranslate"><span class="pre">'a'</span></code>s are matched by the final407<code class="docutils literal notranslate"><span class="pre">aa</span></code> in the pattern.408<code class="docutils literal notranslate"><span class="pre">x{m,n}+</span></code> is equivalent to <code class="docutils literal notranslate"><span class="pre">(?>x{m,n})</span></code>.</p>409<div class="versionadded">410<p><span class="versionmodified added">Added in version 3.11.</span></p>411</div>412</dd>413</dl>414<dl id="index-9">415<dt><code class="docutils literal notranslate"><span class="pre">\</span></code></dt><dd><p>Either escapes special characters (permitting you to match characters like416<code class="docutils literal notranslate"><span class="pre">'*'</span></code>, <code class="docutils literal notranslate"><span class="pre">'?'</span></code>, and so forth), or signals a special sequence; special417sequences are discussed below.</p>418<p>If you’re not using a raw string to express the pattern, remember that Python419also uses the backslash as an escape sequence in string literals; if the escape420sequence isn’t recognized by Python’s parser, the backslash and subsequent421character are included in the resulting string. However, if Python would422recognize the resulting sequence, the backslash should be repeated twice. This423is complicated and hard to understand, so it’s highly recommended that you use424raw strings for all but the simplest expressions.</p>425</dd>426</dl>427<dl id="index-10">428<dt><code class="docutils literal notranslate"><span class="pre">[]</span></code></dt><dd><p>Used to indicate a set of characters. In a set:</p>429<ul class="simple">430<li><p>Characters can be listed individually, e.g. <code class="docutils literal notranslate"><span class="pre">[amk]</span></code> will match <code class="docutils literal notranslate"><span class="pre">'a'</span></code>,431<code class="docutils literal notranslate"><span class="pre">'m'</span></code>, or <code class="docutils literal notranslate"><span class="pre">'k'</span></code>.</p></li>432</ul>433<ul class="simple" id="index-11">434<li><p>Ranges of characters can be indicated by giving two characters and separating435them by a <code class="docutils literal notranslate"><span class="pre">'-'</span></code>, for example <code class="docutils literal notranslate"><span class="pre">[a-z]</span></code> will match any lowercase ASCII letter,436<code class="docutils literal notranslate"><span class="pre">[0-5][0-9]</span></code> will match all the two-digits numbers from <code class="docutils literal notranslate"><span class="pre">00</span></code> to <code class="docutils literal notranslate"><span class="pre">59</span></code>, and437<code class="docutils literal notranslate"><span class="pre">[0-9A-Fa-f]</span></code> will match any hexadecimal digit. If <code class="docutils literal notranslate"><span class="pre">-</span></code> is escaped (e.g.438<code class="docutils literal notranslate"><span class="pre">[a\-z]</span></code>) or if it’s placed as the first or last character439(e.g. <code class="docutils literal notranslate"><span class="pre">[-a]</span></code> or <code class="docutils literal notranslate"><span class="pre">[a-]</span></code>), it will match a literal <code class="docutils literal notranslate"><span class="pre">'-'</span></code>.</p></li>440<li><p>Special characters except backslash lose their special meaning inside sets.441For example,442<code class="docutils literal notranslate"><span class="pre">[(+*)]</span></code> will match any of the literal characters <code class="docutils literal notranslate"><span class="pre">'('</span></code>, <code class="docutils literal notranslate"><span class="pre">'+'</span></code>,443<code class="docutils literal notranslate"><span class="pre">'*'</span></code>, or <code class="docutils literal notranslate"><span class="pre">')'</span></code>.</p></li>444</ul>445<ul class="simple" id="index-12">446<li><p>Backslash either escapes characters which have special meaning in a set447such as <code class="docutils literal notranslate"><span class="pre">'-'</span></code>, <code class="docutils literal notranslate"><span class="pre">']'</span></code>, <code class="docutils literal notranslate"><span class="pre">'^'</span></code> and <code class="docutils literal notranslate"><span class="pre">'\\'</span></code> itself or signals448a special sequence which represents a single character such as449<code class="docutils literal notranslate"><span class="pre">\xa0</span></code> or <code class="docutils literal notranslate"><span class="pre">\n</span></code> or a character class such as <code class="docutils literal notranslate"><span class="pre">\w</span></code> or <code class="docutils literal notranslate"><span class="pre">\S</span></code>450(defined below).451Note that <code class="docutils literal notranslate"><span class="pre">\b</span></code> represents a single “backspace” character,452not a word boundary as outside a set, and numeric escapes453such as <code class="docutils literal notranslate"><span class="pre">\1</span></code> are always octal escapes, not group references.454Special sequences which do not match a single character such as <code class="docutils literal notranslate"><span class="pre">\A</span></code>455and <code class="docutils literal notranslate"><span class="pre">\z</span></code> are not allowed.</p></li>456</ul>457<ul class="simple" id="index-13">458<li><p>Characters that are not within a range can be matched by <em class="dfn">complementing</em>459the set. If the first character of the set is <code class="docutils literal notranslate"><span class="pre">'^'</span></code>, all the characters460that are <em>not</em> in the set will be matched. For example, <code class="docutils literal notranslate"><span class="pre">[^5]</span></code> will match461any character except <code class="docutils literal notranslate"><span class="pre">'5'</span></code>, and <code class="docutils literal notranslate"><span class="pre">[^^]</span></code> will match any character except462<code class="docutils literal notranslate"><span class="pre">'^'</span></code>. <code class="docutils literal notranslate"><span class="pre">^</span></code> has no special meaning if it’s not the first character in463the set.</p></li>464<li><p>To match a literal <code class="docutils literal notranslate"><span class="pre">']'</span></code> inside a set, precede it with a backslash, or465place it at the beginning of the set. For example, both <code class="docutils literal notranslate"><span class="pre">[()[\]{}]</span></code> and466<code class="docutils literal notranslate"><span class="pre">[]()[{}]</span></code> will match a right bracket, as well as left bracket, braces,467and parentheses.</p></li>468</ul>469<ul class="simple">470<li><p>Support of nested sets and set operations as in <a class="reference external" href="https://unicode.org/reports/tr18/">Unicode Technical471Standard #18</a> might be added in the future. This would change the472syntax, so to facilitate this change a <a class="reference internal" href="exceptions.html#FutureWarning" title="FutureWarning"><code class="xref py py-exc docutils literal notranslate"><span class="pre">FutureWarning</span></code></a> will be raised473in ambiguous cases for the time being.474That includes sets starting with a literal <code class="docutils literal notranslate"><span class="pre">'['</span></code> or containing literal475character sequences <code class="docutils literal notranslate"><span class="pre">'--'</span></code>, <code class="docutils literal notranslate"><span class="pre">'&&'</span></code>, <code class="docutils literal notranslate"><span class="pre">'~~'</span></code>, and <code class="docutils literal notranslate"><span class="pre">'||'</span></code>. To476avoid a warning escape them with a backslash.</p></li>477</ul>478<div class="versionchanged">479<p><span class="versionmodified changed">Changed in version 3.7: </span><a class="reference internal" href="exceptions.html#FutureWarning" title="FutureWarning"><code class="xref py py-exc docutils literal notranslate"><span class="pre">FutureWarning</span></code></a> is raised if a character set contains constructs480that will change semantically in the future.</p>481</div>482</dd>483</dl>484<dl class="simple" id="index-14">485<dt><code class="docutils literal notranslate"><span class="pre">|</span></code></dt><dd><p><code class="docutils literal notranslate"><span class="pre">A|B</span></code>, where <em>A</em> and <em>B</em> can be arbitrary REs, creates a regular expression that486will match either <em>A</em> or <em>B</em>. An arbitrary number of REs can be separated by the487<code class="docutils literal notranslate"><span class="pre">'|'</span></code> in this way. This can be used inside groups (see below) as well. As488the target string is scanned, REs separated by <code class="docutils literal notranslate"><span class="pre">'|'</span></code> are tried from left to489right. When one pattern completely matches, that branch is accepted. This means490that once <em>A</em> matches, <em>B</em> will not be tested further, even if it would491produce a longer overall match. In other words, the <code class="docutils literal notranslate"><span class="pre">'|'</span></code> operator is never492greedy. To match a literal <code class="docutils literal notranslate"><span class="pre">'|'</span></code>, use <code class="docutils literal notranslate"><span class="pre">\|</span></code>, or enclose it inside a493character class, as in <code class="docutils literal notranslate"><span class="pre">[|]</span></code>.</p>494</dd>495</dl>496<dl class="simple" id="index-15">497<dt><code class="docutils literal notranslate"><span class="pre">(...)</span></code></dt><dd><p>Matches whatever regular expression is inside the parentheses, and indicates the498start and end of a group; the contents of a group can be retrieved after a match499has been performed, and can be matched later in the string with the <code class="docutils literal notranslate"><span class="pre">\number</span></code>500special sequence, described below. To match the literals <code class="docutils literal notranslate"><span class="pre">'('</span></code> or <code class="docutils literal notranslate"><span class="pre">')'</span></code>,501use <code class="docutils literal notranslate"><span class="pre">\(</span></code> or <code class="docutils literal notranslate"><span class="pre">\)</span></code>, or enclose them inside a character class: <code class="docutils literal notranslate"><span class="pre">[(]</span></code>, <code class="docutils literal notranslate"><span class="pre">[)]</span></code>.</p>502</dd>503</dl>504<dl id="index-16">505<dt><code class="docutils literal notranslate"><span class="pre">(?...)</span></code></dt><dd><p>This is an extension notation (a <code class="docutils literal notranslate"><span class="pre">'?'</span></code> following a <code class="docutils literal notranslate"><span class="pre">'('</span></code> is not meaningful506otherwise). The first character after the <code class="docutils literal notranslate"><span class="pre">'?'</span></code> determines what the meaning507and further syntax of the construct is. Extensions usually do not create a new508group; <code class="docutils literal notranslate"><span class="pre">(?P<name>...)</span></code> is the only exception to this rule. Following are the509currently supported extensions.</p>510</dd>511<dt><code class="docutils literal notranslate"><span class="pre">(?aiLmsux)</span></code></dt><dd><p>(One or more letters from the set512<code class="docutils literal notranslate"><span class="pre">'a'</span></code>, <code class="docutils literal notranslate"><span class="pre">'i'</span></code>, <code class="docutils literal notranslate"><span class="pre">'L'</span></code>, <code class="docutils literal notranslate"><span class="pre">'m'</span></code>, <code class="docutils literal notranslate"><span class="pre">'s'</span></code>, <code class="docutils literal notranslate"><span class="pre">'u'</span></code>, <code class="docutils literal notranslate"><span class="pre">'x'</span></code>.)513The group matches the empty string;514the letters set the corresponding flags for the entire regular expression:</p>515<ul class="simple">516<li><p><a class="reference internal" href="#re.A" title="re.A"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.A</span></code></a> (ASCII-only matching)</p></li>517<li><p><a class="reference internal" href="#re.I" title="re.I"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.I</span></code></a> (ignore case)</p></li>518<li><p><a class="reference internal" href="#re.L" title="re.L"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.L</span></code></a> (locale dependent)</p></li>519<li><p><a class="reference internal" href="#re.M" title="re.M"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.M</span></code></a> (multi-line)</p></li>520<li><p><a class="reference internal" href="#re.S" title="re.S"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.S</span></code></a> (dot matches all)</p></li>521<li><p><a class="reference internal" href="#re.U" title="re.U"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.U</span></code></a> (Unicode matching)</p></li>522<li><p><a class="reference internal" href="#re.X" title="re.X"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.X</span></code></a> (verbose)</p></li>523</ul>524<p>(The flags are described in <a class="reference internal" href="#contents-of-module-re"><span class="std std-ref">Module Contents</span></a>.)525This is useful if you wish to include the flags as part of the526regular expression, instead of passing a <em>flag</em> argument to the527<a class="reference internal" href="#re.compile" title="re.compile"><code class="xref py py-func docutils literal notranslate"><span class="pre">re.compile()</span></code></a> function.528Flags should be used first in the expression string.</p>529<div class="versionchanged">530<p><span class="versionmodified changed">Changed in version 3.11: </span>This construction can only be used at the start of the expression.</p>531</div>532</dd>533</dl>534<dl id="index-17">535<dt><code class="docutils literal notranslate"><span class="pre">(?:...)</span></code></dt><dd><p>A non-capturing version of regular parentheses. Matches whatever regular536expression is inside the parentheses, but the substring matched by the group537<em>cannot</em> be retrieved after performing a match or referenced later in the538pattern.</p>539</dd>540<dt><code class="docutils literal notranslate"><span class="pre">(?aiLmsux-imsx:...)</span></code></dt><dd><p>(Zero or more letters from the set541<code class="docutils literal notranslate"><span class="pre">'a'</span></code>, <code class="docutils literal notranslate"><span class="pre">'i'</span></code>, <code class="docutils literal notranslate"><span class="pre">'L'</span></code>, <code class="docutils literal notranslate"><span class="pre">'m'</span></code>, <code class="docutils literal notranslate"><span class="pre">'s'</span></code>, <code class="docutils literal notranslate"><span class="pre">'u'</span></code>, <code class="docutils literal notranslate"><span class="pre">'x'</span></code>,542optionally followed by <code class="docutils literal notranslate"><span class="pre">'-'</span></code> followed by543one or more letters from the <code class="docutils literal notranslate"><span class="pre">'i'</span></code>, <code class="docutils literal notranslate"><span class="pre">'m'</span></code>, <code class="docutils literal notranslate"><span class="pre">'s'</span></code>, <code class="docutils literal notranslate"><span class="pre">'x'</span></code>.)544The letters set or remove the corresponding flags for the part of the expression:</p>545<ul class="simple">546<li><p><a class="reference internal" href="#re.A" title="re.A"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.A</span></code></a> (ASCII-only matching)</p></li>547<li><p><a class="reference internal" href="#re.I" title="re.I"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.I</span></code></a> (ignore case)</p></li>548<li><p><a class="reference internal" href="#re.L" title="re.L"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.L</span></code></a> (locale dependent)</p></li>549<li><p><a class="reference internal" href="#re.M" title="re.M"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.M</span></code></a> (multi-line)</p></li>550<li><p><a class="reference internal" href="#re.S" title="re.S"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.S</span></code></a> (dot matches all)</p></li>551<li><p><a class="reference internal" href="#re.U" title="re.U"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.U</span></code></a> (Unicode matching)</p></li>552<li><p><a class="reference internal" href="#re.X" title="re.X"><code class="xref py py-const docutils literal notranslate"><span class="pre">re.X</span></code></a> (verbose)</p></li>553</ul>554<p>(The flags are described in <a class="reference internal" href="#contents-of-module-re"><span class="std std-ref">Module Contents</span></a>.)</p>555<p>The letters <code class="docutils literal notranslate"><span class="pre">'a'</span></code>, <code class="docutils literal notranslate"><span class="pre">'L'</span></code> and <code class="docutils literal notranslate"><span class="pre">'u'</span></code> are mutually exclusive when used556as inline flags, so they can’t be combined or follow <code class="docutils literal notranslate"><span class="pre">'-'</span></code>. Instead,557when one of them appears in an inline group, it overrides the matching mode558in the enclosing group. In Unicode patterns <code class="docutils literal notranslate"><span class="pre">(?a:...)</span></code> switches to559ASCII-only matching, and <code class="docutils literal notranslate"><span class="pre">(?u:...)</span></code> switches to Unicode matching560(default). In bytes patterns <code class="docutils literal notranslate"><span class="pre">(?L:...)</span></code> switches to locale dependent561matching, and <code class="docutils literal notranslate"><span class="pre">(?a:...)</span></code> switches to ASCII-only matching (default).562This override is only in effect for the narrow inline group, and the563original matching mode is restored outside of the group.</p>564<div class="versionadded">565<p><span class="versionmodified added">Added in version 3.6.</span></p>566</div>567<div class="versionchanged">568<p><span class="versionmodified changed">Changed in version 3.7: </span>The letters <code class="docutils literal notranslate"><span class="pre">'a'</span></code>, <code class="docutils literal notranslate"><span class="pre">'L'</span></code> and <code class="docutils literal notranslate"><span class="pre">'u'</span></code> also can be used in a group.</p>569</div>570</dd>571<dt><code class="docutils literal notranslate"><span class="pre">(?>...)</span></code></dt><dd><p>Attempts to match <code class="docutils literal notranslate"><span class="pre">...</span></code> as if it was a separate regular expression, and572if successful, continues to match the rest of the pattern following it.573If the subsequent pattern fails to match, the stack can only be unwound574to a point <em>before</em> the <code class="docutils literal notranslate"><span class="pre">(?>...)</span></code> because once exited, the expression,575known as an <em class="dfn">atomic group</em>, has thrown away all stack points within576itself.577Thus, <code class="docutils literal notranslate"><span class="pre">(?>.*).</span></code> would never match anything because first the <code class="docutils literal notranslate"><span class="pre">.*</span></code>578would match all characters possible, then, having nothing left to match,579the final <code class="docutils literal notranslate"><span class="pre">.</span></code> would fail to match.580Since there are no stack points saved in the Atomic Group, and there is581no stack point before it, the entire expression would thus fail to match.</p>582<div class="versionadded">583<p><span class="versionmodified added">Added in version 3.11.</span></p>584</div>585</dd>586</dl>587<dl id="index-18">588<dt><code class="docutils literal notranslate"><span class="pre">(?P<name>...)</span></code></dt><dd><p>Similar to regular parentheses, but the substring matched by the group is589accessible via the symbolic group name <em>name</em>. Group names must be valid590Python identifiers, and in <a class="reference internal" href="stdtypes.html#bytes" title="bytes"><code class="xref py py-class docutils literal notranslate"><span class="pre">bytes</span></code></a> patterns they can only contain591bytes in the ASCII range. Each group name must be defined only once within592a regular expression. A symbolic group is also a numbered group, just as if593the group were not named.</p>594<p>Named groups can be referenced in three contexts. If the pattern is595<code class="docutils literal notranslate"><span class="pre">(?P<quote>['"]).*?(?P=quote)</span></code> (i.e. matching a string quoted with either596single or double quotes):</p>597<table class="docutils align-default">598<thead>599<tr class="row-odd"><th class="head"><p>Context of reference to group “quote”</p></th>600<th class="head"><p>Ways to reference it</p></th>601</tr>602</thead>603<tbody>604<tr class="row-even"><td><p>in the same pattern itself</p></td>605<td><ul class="simple">606<li><p><code class="docutils literal notranslate"><span class="pre">(?P=quote)</span></code> (as shown)</p></li>607<li><p><code class="docutils literal notranslate"><span class="pre">\1</span></code></p></li>608</ul>609</td>610</tr>611<tr class="row-odd"><td><p>when processing match object <em>m</em></p></td>612<td><ul class="simple">613<li><p><code class="docutils literal notranslate"><span class="pre">m.group('quote')</span></code></p></li>614<li><p><code class="docutils literal notranslate"><span class="pre">m.end('quote')</span></code> (etc.)</p></li>615</ul>616</td>617</tr>618<tr class="row-even"><td><p>in a string passed to the <em>repl</em>619argument of <code class="docutils literal notranslate"><span class="pre">re.sub()</span></code></p></td>620<td><ul class="simple">621<li><p><code class="docutils literal notranslate"><span class="pre">\g<quote></span></code></p></li>622<li><p><code class="docutils literal notranslate"><span class="pre">\g<1></span></code></p></li>623<li><p><code class="docutils literal notranslate"><span class="pre">\1</span></code></p></li>624</ul>625</td>626</tr>627</tbody>628</table>629<div class="versionchanged">630<p><span class="versionmodified changed">Changed in version 3.12: </span>In <a class="reference internal" href="stdtypes.html#bytes" title="bytes"><code class="xref py py-class docutils literal notranslate"><span class="pre">bytes</span></code></a> patterns, group <em>name</em> can only contain bytes631in the ASCII range (<code class="docutils literal notranslate"><span class="pre">b'\x00'</span></code>-<code class="docutils literal notranslate"><span class="pre">b'\x7f'</span></code>).</p>632</div>633</dd>634</dl>635<dl class="simple" id="index-19">636<dt><code class="docutils literal notranslate"><span class="pre">(?P=name)</span></code></dt><dd><p>A backreference to a named group; it matches whatever text was matched by the637earlier group named <em>name</em>.</p>638</dd>639</dl>640<dl class="simple" id="index-20">641<dt><code class="docutils literal notranslate"><span class="pre">(?#...)</span></code></dt><dd><p>A comment; the contents of the parentheses are simply ignored.</p>642</dd>643</dl>644<dl class="simple" id="index-21">645<dt><code class="docutils literal notranslate"><span class="pre">(?=...)</span></code></dt><dd><p>Matches if <code class="docutils literal notranslate"><span class="pre">...</span></code> matches next, but doesn’t consume any of the string. This is646called a <em class="dfn">lookahead assertion</em>. For example, <code class="docutils literal notranslate"><span class="pre">Isaac</span> <span class="pre">(?=Asimov)</span></code> will match647<code class="docutils literal notranslate"><span class="pre">'Isaac</span> <span class="pre">'</span></code> only if it’s followed by <code class="docutils literal notranslate"><span class="pre">'Asimov'</span></code>.</p>648</dd>649</dl>650<dl class="simple" id="index-22">651<dt><code class="docutils literal notranslate"><span class="pre">(?!...)</span></code></dt><dd><p>Matches if <code class="docutils literal notranslate"><span class="pre">...</span></code> doesn’t match next. This is a <em class="dfn">negative lookahead assertion</em>.652For example, <code class="docutils literal notranslate"><span class="pre">Isaac</span> <span class="pre">(?!Asimov)</span></code> will match <code class="docutils literal notranslate"><span class="pre">'Isaac</span> <span class="pre">'</span></code> only if it’s <em>not</em>653followed by <code class="docutils literal notranslate"><span class="pre">'Asimov'</span></code>.</p>654</dd>655</dl>656<dl id="index-23">657<dt><code class="docutils literal notranslate"><span class="pre">(?<=...)</span></code></dt><dd><p>Matches if the current position in the string is preceded by a match for <code class="docutils literal notranslate"><span class="pre">...</span></code>658that ends at the current position. This is called a <em class="dfn">positive lookbehind659assertion</em>. <code class="docutils literal notranslate"><span class="pre">(?<=abc)def</span></code> will find a match in <code class="docutils literal notranslate"><span class="pre">'abcdef'</span></code>, since the660lookbehind will back up 3 characters and check if the contained pattern matches.661The contained pattern must only match strings of some fixed length, meaning that662<code class="docutils literal notranslate"><span class="pre">abc</span></code> or <code class="docutils literal notranslate"><span class="pre">a|b</span></code> are allowed, but <code class="docutils literal notranslate"><span class="pre">a*</span></code> and <code class="docutils literal notranslate"><span class="pre">a{3,4}</span></code> are not. Note that663patterns which start with positive lookbehind assertions will not match at the664beginning of the string being searched; you will most likely want to use the665<a class="reference internal" href="#re.search" title="re.search"><code class="xref py py-func docutils literal notranslate"><span class="pre">search()</span></code></a> function rather than the <a class="reference internal" href="#re.match" title="re.match"><code class="xref py py-func docutils literal notranslate"><span class="pre">match()</span></code></a> function:</p>666<div class="doctest highlight-default notranslate"><div class="highlight"><pre><span></span><span class="gp">>>> </span><span class="kn">import</span><span class="w"> </span><span class="nn">re</span>667<span class="gp">>>> </span><span class="n">m</span> <span class="o">=</span> <span class="n">re</span><span class="o">.</span><span class="n">search</span><span class="p">(</span><span class="s1">'(?<=abc)def'</span><span class="p">,</span> <span class="s1">'abcdef'</span><span class="p">)</span>668<span class="gp">>>> </span><span class="n">m</span><span class="o">.</span><span class="n">group</span><span class="p">(</span><span class="mi">0</span><span class="p">)</span>669<span class="go">'def'</span>670</pre></div>671</div>672<p>This example looks for a word following a hyphen:</p>673<div class="doctest highlight-default notranslate"><div class="highlight"><pre><span></span><span class="gp">>>> </span><span class="n">m</span> <span class="o">=</span> <span class="n">re</span><span class="o">.</span><span class="n">search</span><span class="p">(</span><span class="sa">r</span><span class="s1">'(?<=-)\w+'</span><span class="p">,</span> <span class="s1">'spam-egg'</span><span class="p">)</span>674<span class="gp">>>> </span><span class="n">m</span><span class="o">.</span><span class="n">group</span><span class="p">(</span><span class="mi">0</span><span class="p">)</span>675<span class="go">'egg'</span>676</pre></div>677</div>678<div class="versionchanged">679<p><span class="versionmodified changed">Changed in version 3.5: </span>Added support for group references of fixed length.</p>680</div>681</dd>682</dl>683<dl class="simple" id="index-24">684<dt><code class="docutils literal notranslate"><span class="pre">(?<!...)</span></code></dt><dd><p>Matches if the current position in the string is not preceded by a match for685<code class="docutils literal notranslate"><span class="pre">...</span></code>. This is called a <em class="dfn">negative lookbehind assertion</em>. Similar to686positive lookbehind assertions, the contained pattern must only match strings of687some fixed length. Patterns which start with negative lookbehind assertions may688match at the beginning of the string being searched.</p>689</dd>690</dl>691<span id="re-conditional-expression"></span><dl id="index-25">692<dt><code class="docutils literal notranslate"><span class="pre">(?(id/name)yes-pattern|no-pattern)</span></code></dt><dd><p>Will try to match with <code class="docutils literal notranslate"><span class="pre">yes-pattern</span></code> if the group with given <em>id</em> or693<em>name</em> exists, and with <code class="docutils literal notranslate"><span class="pre">no-pattern</span></code> if it doesn’t. <code class="docutils literal notranslate"><span class="pre">no-pattern</span></code> is694optional and can be omitted. For example,695<code class="docutils literal notranslate"><span class="pre">(<)?(\w+@\w+(?:\.\w+)+)(?(1)>|$)</span></code> is a poor email matching pattern, which696will match with <code class="docutils literal notranslate"><span class="pre">'<user@host.com>'</span></code> as well as <code class="docutils literal notranslate"><span class="pre">'user@host.com'</span></code>, but697not with <code class="docutils literal notranslate"><span class="pre">'<user@host.com'</span></code> nor <code class="docutils literal notranslate"><span class="pre">'user@host.com>'</span></code>.</p>698<div class="versionchanged">699<p><span class="versionmodified changed">Changed in version 3.12: </span>Group <em>id</em> can only contain ASCII digits.700In <a class="reference internal" href="stdtypes.html#bytes" title="bytes"><code class="xref py py-class docutils literal notranslate"><span class="pre">bytes</span></code></a> patterns, group <em>name</em> can only contain bytes701in the ASCII range (<code class="docutils literal notranslate"><span class="pre">b'\x00'</span></code>-<code class="docutils literal notranslate"><span class="pre">b'\x7f'</span></code>).</p>702</div>703</dd>704</dl>705<p id="re-special-sequences">The special sequences consist of <code class="docutils literal notranslate"><span class="pre">'\'</span></code> and a character from the list below.706If the ordinary character is not an ASCII digit or an ASCII letter, then the707resulting RE will match the second character. For example, <code class="docutils literal notranslate"><span class="pre">\$</span></code> matches the708character <code class="docutils literal notranslate"><span class="pre">'$'</span></code>.</p>709<dl class="simple" id="index-26">710<dt><code class="docutils literal notranslate"><span class="pre">\number</span></code></dt><dd><p>Matches the contents of the group of the same number. Groups are numbered711starting from 1. For example, <code class="docutils literal notranslate"><span class="pre">(.+)</span> <span class="pre">\1</span></code> matches <code class="docutils literal notranslate"><span class="pre">'the</span> <span class="pre">the'</span></code> or <code class="docutils literal notranslate"><span class="pre">'55</span> <span class="pre">55'</span></code>,712but not <code class="docutils literal notranslate"><span class="pre">'thethe'</span></code> (note the space after the group). This special sequence713can only be used to match one of the first 99 groups. If the first digit of714<em>number</em> is 0, or <em>number</em> is 3 octal digits long, it will not be interpreted as715a group match, but as the character with octal value <em>number</em>. Inside the716<code class="docutils literal notranslate"><span class="pre">'['</span></code> and <code class="docutils literal notranslate"><span class="pre">']'</span></code> of a character class, all numeric escapes are treated as717characters.</p>718</dd>719</dl>720<dl class="simple" id="index-27">721<dt><code class="docutils literal notranslate"><span class="pre">\A</span></code></dt><dd><p>Matches only at the start of the string.</p>722</dd>723</dl>724<dl id="index-28">725<dt><code class="docutils literal notranslate"><span class="pre">\b</span></code></dt><dd><p>Matches the empty string, but only at the beginning or end of a word.726A word is defined as a sequence of word characters.727Note that formally, <code class="docutils literal notranslate"><span class="pre">\b</span></code> is defined as the boundary728between a <code class="docutils literal notranslate"><span class="pre">\w</span></code> and a <code class="docutils literal notranslate"><span class="pre">\W</span></code> character (or vice versa),729or between <code class="docutils literal notranslate"><span class="pre">\w</span></code> and the beginning or end of the string.730This means that <code class="docutils literal notranslate"><span class="pre">r'\bat\b'</span></code> matches <code class="docutils literal notranslate"><span class="pre">'at'</span></code>, <code class="docutils literal notranslate"><span class="pre">'at.'</span></code>, <code class="docutils literal notranslate"><span class="pre">'(at)'</span></code>,731and <code class="docutils literal notranslate"><span class="pre">'as</span> <span class="pre">at</span> <span class="pre">ay'</span></code> but not <code class="docutils literal notranslate"><span class="pre">'attempt'</span></code> or <code class="docutils literal notranslate"><span class="pre">'atlas'</span></code>.</p>732<p>The default word characters in Unicode (str) patterns733are Unicode alphanumerics and the underscore,734but this can be changed by using the <a class="reference internal" href="#re.ASCII" title="re.ASCII"><code class="xref py py-const docutils literal notranslate"><span class="pre">ASCII</span></code></a> flag.735Word boundaries are determined by the current locale736if the <a class="reference internal" href="#re.LOCALE" title="re.LOCALE"><code class="xref py py-const docutils literal notranslate"><span class="pre">LOCALE</span></code></a> flag is used.</p>737<div class="admonition note">738<p class="admonition-title">Note</p>739<p>Inside a character range, <code class="docutils literal notranslate"><span class="pre">\b</span></code> represents the backspace character,740for compatibility with Python’s string literals.</p>741</div>742</dd>743</dl>744<dl id="index-29">745<dt><code class="docutils literal notranslate"><span class="pre">\B</span></code></dt><dd><p>Matches the empty string,746but only when it is <em>not</em> at the beginning or end of a word.747This means that <code class="docutils literal notranslate"><span class="pre">r'at\B'</span></code> matches <code class="docutils literal notranslate"><span class="pre">'athens'</span></code>, <code class="docutils literal notranslate"><span class="pre">'atom'</span></code>,748<code class="docutils literal notranslate"><span class="pre">'attorney'</span></code>, but not <code class="docutils literal notranslate"><span class="pre">'at'</span></code>, <code class="docutils literal notranslate"><span class="pre">'at.'</span></code>, or <code class="docutils literal notranslate"><span class="pre">'at!'</span></code>.749<code class="docutils literal notranslate"><span class="pre">\B</span></code> is the opposite of <code class="docutils literal notranslate"><span class="pre">\b</span></code>,750so word characters in Unicode (str) patterns751are Unicode alphanumerics or the underscore,752although this can be changed by using the <a class="reference internal" href="#re.ASCII" title="re.ASCII"><code class="xref py py-const docutils literal notranslate"><span class="pre">ASCII</span></code></a> flag.753Word boundaries are determined by the current locale754if the <a class="reference internal" href="#re.LOCALE" title="re.LOCALE"><code class="xref py py-const docutils literal notranslate"><span class="pre">LOCALE</span></code></a> flag is used.</p>755<div class="versionchanged">756<p><span class="versionmodified changed">Changed in version 3.14: </span><code class="docutils literal notranslate"><span class="pre">\B</span></code> now matches empty input string.</p>757</div>758</dd>759</dl>760<dl id="index-30">761<dt><code class="docutils literal notranslate"><span class="pre">\d</span></code></dt><dd><dl>762<dt>For Unicode (str) patterns:</dt><dd><p>Matches any Unicode decimal digit763(that is, any character in Unicode character category <a class="reference external" href="https://www.unicode.org/versions/Unicode15.0.0/ch04.pdf#G134153">[Nd]</a>).764This includes <code class="docutils literal notranslate"><span class="pre">[0-9]</span></code>, and also many other digit characters.</p>765<p>Matches <code class="docutils literal notranslate"><span class="pre">[0-9]</span></code> if the <a class="reference internal" href="#re.ASCII" title="re.ASCII"><code class="xref py py-const docutils literal notranslate"><span class="pre">ASCII</span></code></a> flag is used.</p>766</dd>767<dt>For 8-bit (bytes) patterns:</dt><dd><p>Matches any decimal digit in the ASCII character set;768this is equivalent to <code class="docutils literal notranslate"><span class="pre">[0-9]</span></code>.</p>769</dd>770</dl>771</dd>772</dl>773<dl id="index-31">774<dt><code class="docutils literal notranslate"><span class="pre">\D</span></code></dt><dd><p>Matches any character which is not a decimal digit.775This is the opposite of <code class="docutils literal notranslate"><span class="pre">\d</span></code>.</p>776<p>Matches <code class="docutils literal notranslate"><span class="pre">[^0-9]</span></code> if the <a class="reference internal" href="#re.ASCII" title="re.ASCII"><code class="xref py py-const docutils literal notranslate"><span class="pre">ASCII</span></code></a> flag is used.</p>777</dd>778</dl>779<dl id="index-32">780<dt><code class="docutils literal notranslate"><span class="pre">\s</span></code></dt><dd><dl>781<dt>For Unicode (str) patterns:</dt><dd><p>Matches Unicode whitespace characters (as defined by <a class="reference internal" href="stdtypes.html#str.isspace" title="str.isspace"><code class="xref py py-meth docutils literal notranslate"><span class="pre">str.isspace()</span></code></a>).782This includes <code class="docutils literal notranslate"><span class="pre">[</span> <span class="pre">\t\n\r\f\v]</span></code>, and also many other characters, for example the783non-breaking spaces mandated by typography rules in many languages.</p>784<p>Matches <code class="docutils literal notranslate"><span class="pre">[</span> <span class="pre">\t\n\r\f\v]</span></code> if the <a class="reference internal" href="#re.ASCII" title="re.ASCII"><code class="xref py py-const docutils literal notranslate"><span class="pre">ASCII</span></code></a> flag is used.</p>785</dd>786<dt>For 8-bit (bytes) patterns:</dt><dd><p>Matches characters considered whitespace in the ASCII character set;787this is equivalent to <code class="docutils literal notranslate"><span class="pre">[</span> <span class="pre">\t\n\r\f\v]</span></code>.</p>788</dd>789</dl>790</dd>791</dl>792<dl id="index-33">793<dt><code class="docutils literal notranslate"><span class="pre">\S</span></code></dt><dd><p>Matches any character which is not a whitespace character. This is794the opposite of <code class="docutils literal notranslate"><span class="pre">\s</span></code>.</p>795<p>Matches <code class="docutils literal notranslate"><span class="pre">[^</span> <span class="pre">\t\n\r\f\v]</span></code> if the <a class="reference internal" href="#re.ASCII" title="re.ASCII"><code class="xref py py-const docutils literal notranslate"><span class="pre">ASCII</span></code></a> flag is used.</p>796</dd>797</dl>798<dl id="index-34">799<dt><code class="docutils literal notranslate"><span class="pre">\w</span></code></dt><dd><dl>800<dt>For Unicode (str) patterns:</dt><dd><p>Matches Unicode word characters;801this includes all Unicode alphanumeric characters802(as defined by <a class="reference internal" href="stdtypes.html#str.isalnum" title="str.isalnum"><code class="xref py py-meth docutils literal notranslate"><span class="pre">str.isalnum()</span></code></a>),803as well as the underscore (<code class="docutils literal notranslate"><span class="pre">_</span></code>).</p>804<p>Matches <code class="docutils literal notranslate"><span class="pre">[a-zA-Z0-9_]</span></code> if the <a class="reference internal" href="#re.ASCII" title="re.ASCII"><code class="xref py py-const docutils literal notranslate"><span class="pre">ASCII</span></code></a> flag is used.</p>805</dd>806<dt>For 8-bit (bytes) patterns:</dt><dd><p>Matches characters considered alphanumeric in the ASCII character set;807this is equivalent to <code class="docutils literal notranslate"><span class="pre">[a-zA-Z0-9_]</span></code>.808If the <a class="reference internal" href="#re.LOCALE" title="re.LOCALE"><code class="xref py py-const docutils literal notranslate"><span class="pre">LOCALE</span></code></a> flag is used,809matches characters considered alphanumeric in the current locale and the underscore.</p>810</dd>811</dl>812</dd>813</dl>814<dl id="index-35">815<dt><code class="docutils literal notranslate"><span class="pre">\W</span></code></dt><dd><p>Matches any character which is not a word character.816This is the opposite of <code class="docutils literal notranslate"><span class="pre">\w</span></code>.817By default, matches non-underscore (<code class="docutils literal notranslate"><span class="pre">_</span></code>) characters818for which <a class="reference internal" href="stdtypes.html#str.isalnum" title="str.isalnum"><code class="xref py py-meth docutils literal notranslate"><span class="pre">str.isalnum()</span></code></a> returns <code class="docutils literal notranslate"><span class="pre">False</span></code>.</p>819<p>Matches <code class="docutils literal notranslate"><span class="pre">[^a-zA-Z0-9_]</span></code> if the <a class="reference internal" href="#re.ASCII" title="re.ASCII"><code class="xref py py-const docutils literal notranslate"><span class="pre">ASCII</span></code></a> flag is used.</p>820<p>If the <a class="reference internal" href="#re.LOCALE" title="re.LOCALE"><code class="xref py py-const docutils literal notranslate"><span class="pre">LOCALE</span></code></a> flag is used,821matches characters which are neither alphanumeric in the current locale822nor the underscore.</p>823</dd>824</dl>825<dl id="index-36">826<dt><code class="docutils literal notranslate"><span class="pre">\z</span></code></dt><dd><p>Matches only at the end of the string.</p>827<div class="versionadded">828<p><span class="versionmodified added">Added in version 3.14.</span></p>829</div>830</dd>831<dt><code class="docutils literal notranslate"><span class="pre">\Z</span></code></dt><dd><p>The same as <code class="docutils literal notranslate"><span class="pre">\z</span></code>. For compatibility with old Python versions.</p>832</dd>833</dl>834<p id="index-37">Most of the <a class="reference internal" href="../reference/lexical_analysis.html#escape-sequences"><span class="std std-ref">escape sequences</span></a> supported by Python835string literals are also accepted by the regular expression parser:</p>836<div class="highlight-python3 notranslate"><div class="highlight"><pre><span></span>\<span class="n">a</span> \<span class="n">b</span> \<span class="n">f</span> \<span class="n">n</span>837\<span class="n">N</span> \<span class="n">r</span> \<span class="n">t</span> \<span class="n">u</span>838\<span class="n">U</span> \<span class="n">v</span> \<span class="n">x</span> \\839</pre></div>840</div>841<p>(Note that <code class="docutils literal notranslate"><span class="pre">\b</span></code> is used to represent word boundaries, and means “backspace”842only inside character classes.)</p>843<p><code class="docutils literal notranslate"><span class="pre">'\u'</span></code>, <code class="docutils literal notranslate"><span class="pre">'\U'</span></code>, and <code class="docutils literal notranslate"><span class="pre">'\N'</span></code> escape sequences are844only recognized in Unicode (str) patterns.845In bytes patterns they are errors.846Unknown escapes of ASCII letters are reserved847for future use and treated as errors.</p>848<p>Octal escapes are included in a limited form. If the first digit is a 0, or if849there are three octal digits, it is considered an octal escape. Otherwise, it is850a group reference. As for string literals, octal escapes are always at most851three digits in length.</p>852<div class="versionchanged">853<p><span class="versionmodified changed">Changed in version 3.3: </span>The <code class="docutils literal notranslate"><span class="pre">'\u'</span></code> and <code class="docutils literal notranslate"><span class="pre">'\U'</span></code> escape sequences have been added.</p>854</div>855<div class="versionchanged">856<p><span class="versionmodified changed">Changed in version 3.6: </span>Unknown escapes consisting of <code class="docutils literal notranslate"><span class="pre">'\'</span></code> and an ASCII letter now are errors.</p>857</div>858<div class="versionchanged">859<p><span class="versionmodified changed">Changed in version 3.8: </span>The <code class="samp docutils literal notranslate"><span class="pre">'\N{</span><em><span class="pre">name</span></em><span class="pre">}'</span></code> escape sequence has been added. As in string literals,860it expands to the named Unicode character (e.g. <code class="docutils literal notranslate"><span class="pre">'\N{EM</span> <span class="pre">DASH}'</span></code>).</p>861</div>862</section>863<section id="module-contents">864<span id="contents-of-module-re"></span><h2>Module Contents<a class="headerlink" href="#module-contents" title="Link to this heading">¶</a></h2>865<p>The module defines several functions, constants, and an exception. Some of the866functions are simplified versions of the full featured methods for compiled867regular expressions. Most non-trivial applications always use the compiled868form.</p>869<section id="flags">870<h3>Flags<a class="headerlink" href="#flags" title="Link to this heading">¶</a></h3>871<div class="versionchanged">872<p><span class="versionmodified changed">Changed in version 3.6: </span>Flag constants are now instances of <a class="reference internal" href="#re.RegexFlag" title="re.RegexFlag"><code class="xref py py-class docutils literal notranslate"><span class="pre">RegexFlag</span></code></a>, which is a subclass of873<a class="reference internal" href="enum.html#enum.IntFlag" title="enum.IntFlag"><code class="xref py py-class docutils literal notranslate"><span class="pre">enum.IntFlag</span></code></a>.</p>874</div>875<dl class="py class">876<dt class="sig sig-object py" id="re.RegexFlag">877<em class="property"><span class="k"><span class="pre">class</span></span><span class="w"> </span></em><span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">RegexFlag</span></span><a class="headerlink" href="#re.RegexFlag" title="Link to this definition">¶</a></dt>878<dd><p>An <a class="reference internal" href="enum.html#enum.IntFlag" title="enum.IntFlag"><code class="xref py py-class docutils literal notranslate"><span class="pre">enum.IntFlag</span></code></a> class containing the regex options listed below.</p>879<div class="versionadded">880<p><span class="versionmodified added">Added in version 3.11: </span>- added to <code class="docutils literal notranslate"><span class="pre">__all__</span></code></p>881</div>882</dd></dl>883 884<dl class="py data">885<dt class="sig sig-object py" id="re.A">886<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">A</span></span><a class="headerlink" href="#re.A" title="Link to this definition">¶</a></dt>887<dt class="sig sig-object py" id="re.ASCII">888<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">ASCII</span></span><a class="headerlink" href="#re.ASCII" title="Link to this definition">¶</a></dt>889<dd><p>Make <code class="docutils literal notranslate"><span class="pre">\w</span></code>, <code class="docutils literal notranslate"><span class="pre">\W</span></code>, <code class="docutils literal notranslate"><span class="pre">\b</span></code>, <code class="docutils literal notranslate"><span class="pre">\B</span></code>, <code class="docutils literal notranslate"><span class="pre">\d</span></code>, <code class="docutils literal notranslate"><span class="pre">\D</span></code>, <code class="docutils literal notranslate"><span class="pre">\s</span></code> and <code class="docutils literal notranslate"><span class="pre">\S</span></code>890perform ASCII-only matching instead of full Unicode matching. This is only891meaningful for Unicode (str) patterns, and is ignored for bytes patterns.</p>892<p>Corresponds to the inline flag <code class="docutils literal notranslate"><span class="pre">(?a)</span></code>.</p>893<div class="admonition note">894<p class="admonition-title">Note</p>895<p>The <a class="reference internal" href="#re.U" title="re.U"><code class="xref py py-const docutils literal notranslate"><span class="pre">U</span></code></a> flag still exists for backward compatibility,896but is redundant in Python 3 since897matches are Unicode by default for <code class="docutils literal notranslate"><span class="pre">str</span></code> patterns,898and Unicode matching isn’t allowed for bytes patterns.899<a class="reference internal" href="#re.UNICODE" title="re.UNICODE"><code class="xref py py-const docutils literal notranslate"><span class="pre">UNICODE</span></code></a> and the inline flag <code class="docutils literal notranslate"><span class="pre">(?u)</span></code> are similarly redundant.</p>900</div>901</dd></dl>902 903<dl class="py data">904<dt class="sig sig-object py" id="re.DEBUG">905<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">DEBUG</span></span><a class="headerlink" href="#re.DEBUG" title="Link to this definition">¶</a></dt>906<dd><p>Display debug information about compiled expression.</p>907<p>No corresponding inline flag.</p>908</dd></dl>909 910<dl class="py data">911<dt class="sig sig-object py" id="re.I">912<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">I</span></span><a class="headerlink" href="#re.I" title="Link to this definition">¶</a></dt>913<dt class="sig sig-object py" id="re.IGNORECASE">914<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">IGNORECASE</span></span><a class="headerlink" href="#re.IGNORECASE" title="Link to this definition">¶</a></dt>915<dd><p>Perform case-insensitive matching;916expressions like <code class="docutils literal notranslate"><span class="pre">[A-Z]</span></code> will also match lowercase letters.917Full Unicode matching (such as <code class="docutils literal notranslate"><span class="pre">Ü</span></code> matching <code class="docutils literal notranslate"><span class="pre">ü</span></code>)918also works unless the <a class="reference internal" href="#re.ASCII" title="re.ASCII"><code class="xref py py-const docutils literal notranslate"><span class="pre">ASCII</span></code></a> flag919is used to disable non-ASCII matches.920The current locale does not change the effect of this flag921unless the <a class="reference internal" href="#re.LOCALE" title="re.LOCALE"><code class="xref py py-const docutils literal notranslate"><span class="pre">LOCALE</span></code></a> flag is also used.</p>922<p>Corresponds to the inline flag <code class="docutils literal notranslate"><span class="pre">(?i)</span></code>.</p>923<p>Note that when the Unicode patterns <code class="docutils literal notranslate"><span class="pre">[a-z]</span></code> or <code class="docutils literal notranslate"><span class="pre">[A-Z]</span></code> are used in924combination with the <a class="reference internal" href="#re.IGNORECASE" title="re.IGNORECASE"><code class="xref py py-const docutils literal notranslate"><span class="pre">IGNORECASE</span></code></a> flag, they will match the 52 ASCII925letters and 4 additional non-ASCII letters: ‘İ’ (U+0130, Latin capital926letter I with dot above), ‘ı’ (U+0131, Latin small letter dotless i),927‘ſ’ (U+017F, Latin small letter long s) and ‘K’ (U+212A, Kelvin sign).928If the <a class="reference internal" href="#re.ASCII" title="re.ASCII"><code class="xref py py-const docutils literal notranslate"><span class="pre">ASCII</span></code></a> flag is used, only letters ‘a’ to ‘z’929and ‘A’ to ‘Z’ are matched.</p>930</dd></dl>931 932<dl class="py data">933<dt class="sig sig-object py" id="re.L">934<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">L</span></span><a class="headerlink" href="#re.L" title="Link to this definition">¶</a></dt>935<dt class="sig sig-object py" id="re.LOCALE">936<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">LOCALE</span></span><a class="headerlink" href="#re.LOCALE" title="Link to this definition">¶</a></dt>937<dd><p>Make <code class="docutils literal notranslate"><span class="pre">\w</span></code>, <code class="docutils literal notranslate"><span class="pre">\W</span></code>, <code class="docutils literal notranslate"><span class="pre">\b</span></code>, <code class="docutils literal notranslate"><span class="pre">\B</span></code> and case-insensitive matching938dependent on the current locale.939This flag can be used only with bytes patterns.</p>940<p>Corresponds to the inline flag <code class="docutils literal notranslate"><span class="pre">(?L)</span></code>.</p>941<div class="admonition warning">942<p class="admonition-title">Warning</p>943<p>This flag is discouraged; consider Unicode matching instead.944The locale mechanism is very unreliable945as it only handles one “culture” at a time946and only works with 8-bit locales.947Unicode matching is enabled by default for Unicode (str) patterns948and it is able to handle different locales and languages.</p>949</div>950<div class="versionchanged">951<p><span class="versionmodified changed">Changed in version 3.6: </span><a class="reference internal" href="#re.LOCALE" title="re.LOCALE"><code class="xref py py-const docutils literal notranslate"><span class="pre">LOCALE</span></code></a> can be used only with bytes patterns952and is not compatible with <a class="reference internal" href="#re.ASCII" title="re.ASCII"><code class="xref py py-const docutils literal notranslate"><span class="pre">ASCII</span></code></a>.</p>953</div>954<div class="versionchanged">955<p><span class="versionmodified changed">Changed in version 3.7: </span>Compiled regular expression objects with the <a class="reference internal" href="#re.LOCALE" title="re.LOCALE"><code class="xref py py-const docutils literal notranslate"><span class="pre">LOCALE</span></code></a> flag956no longer depend on the locale at compile time.957Only the locale at matching time affects the result of matching.</p>958</div>959</dd></dl>960 961<dl class="py data">962<dt class="sig sig-object py" id="re.M">963<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">M</span></span><a class="headerlink" href="#re.M" title="Link to this definition">¶</a></dt>964<dt class="sig sig-object py" id="re.MULTILINE">965<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">MULTILINE</span></span><a class="headerlink" href="#re.MULTILINE" title="Link to this definition">¶</a></dt>966<dd><p>When specified, the pattern character <code class="docutils literal notranslate"><span class="pre">'^'</span></code> matches at the beginning of the967string and at the beginning of each line (immediately following each newline);968and the pattern character <code class="docutils literal notranslate"><span class="pre">'$'</span></code> matches at the end of the string and at the969end of each line (immediately preceding each newline). By default, <code class="docutils literal notranslate"><span class="pre">'^'</span></code>970matches only at the beginning of the string, and <code class="docutils literal notranslate"><span class="pre">'$'</span></code> only at the end of the971string and immediately before the newline (if any) at the end of the string.</p>972<p>Corresponds to the inline flag <code class="docutils literal notranslate"><span class="pre">(?m)</span></code>.</p>973</dd></dl>974 975<dl class="py data">976<dt class="sig sig-object py" id="re.NOFLAG">977<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">NOFLAG</span></span><a class="headerlink" href="#re.NOFLAG" title="Link to this definition">¶</a></dt>978<dd><p>Indicates no flag being applied, the value is <code class="docutils literal notranslate"><span class="pre">0</span></code>. This flag may be used979as a default value for a function keyword argument or as a base value that980will be conditionally ORed with other flags. Example of use as a default981value:</p>982<div class="highlight-python3 notranslate"><div class="highlight"><pre><span></span><span class="k">def</span><span class="w"> </span><span class="nf">myfunc</span><span class="p">(</span><span class="n">text</span><span class="p">,</span> <span class="n">flag</span><span class="o">=</span><span class="n">re</span><span class="o">.</span><span class="n">NOFLAG</span><span class="p">):</span>983 <span class="k">return</span> <span class="n">re</span><span class="o">.</span><span class="n">search</span><span class="p">(</span><span class="n">text</span><span class="p">,</span> <span class="n">flag</span><span class="p">)</span>984</pre></div>985</div>986<div class="versionadded">987<p><span class="versionmodified added">Added in version 3.11.</span></p>988</div>989</dd></dl>990 991<dl class="py data">992<dt class="sig sig-object py" id="re.S">993<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">S</span></span><a class="headerlink" href="#re.S" title="Link to this definition">¶</a></dt>994<dt class="sig sig-object py" id="re.DOTALL">995<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">DOTALL</span></span><a class="headerlink" href="#re.DOTALL" title="Link to this definition">¶</a></dt>996<dd><p>Make the <code class="docutils literal notranslate"><span class="pre">'.'</span></code> special character match any character at all, including a997newline; without this flag, <code class="docutils literal notranslate"><span class="pre">'.'</span></code> will match anything <em>except</em> a newline.</p>998<p>Corresponds to the inline flag <code class="docutils literal notranslate"><span class="pre">(?s)</span></code>.</p>999</dd></dl>1000 1001<dl class="py data">1002<dt class="sig sig-object py" id="re.U">1003<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">U</span></span><a class="headerlink" href="#re.U" title="Link to this definition">¶</a></dt>1004<dt class="sig sig-object py" id="re.UNICODE">1005<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">UNICODE</span></span><a class="headerlink" href="#re.UNICODE" title="Link to this definition">¶</a></dt>1006<dd><p>In Python 3, Unicode characters are matched by default1007for <code class="docutils literal notranslate"><span class="pre">str</span></code> patterns.1008This flag is therefore redundant with <strong>no effect</strong>1009and is only kept for backward compatibility.</p>1010<p>See <a class="reference internal" href="#re.ASCII" title="re.ASCII"><code class="xref py py-const docutils literal notranslate"><span class="pre">ASCII</span></code></a> to restrict matching to ASCII characters instead.</p>1011</dd></dl>1012 1013<dl class="py data">1014<dt class="sig sig-object py" id="re.X">1015<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">X</span></span><a class="headerlink" href="#re.X" title="Link to this definition">¶</a></dt>1016<dt class="sig sig-object py" id="re.VERBOSE">1017<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">VERBOSE</span></span><a class="headerlink" href="#re.VERBOSE" title="Link to this definition">¶</a></dt>1018<dd><p id="index-38">This flag allows you to write regular expressions that look nicer and are1019more readable by allowing you to visually separate logical sections of the1020pattern and add comments. Whitespace within the pattern is ignored, except1021when in a character class, or when preceded by an unescaped backslash,1022or within tokens like <code class="docutils literal notranslate"><span class="pre">*?</span></code>, <code class="docutils literal notranslate"><span class="pre">(?:</span></code> or <code class="docutils literal notranslate"><span class="pre">(?P<...></span></code>. For example, <code class="docutils literal notranslate"><span class="pre">(?</span> <span class="pre">:</span></code>1023and <code class="docutils literal notranslate"><span class="pre">*</span> <span class="pre">?</span></code> are not allowed.1024When a line contains a <code class="docutils literal notranslate"><span class="pre">#</span></code> that is not in a character class and is not1025preceded by an unescaped backslash, all characters from the leftmost such1026<code class="docutils literal notranslate"><span class="pre">#</span></code> through the end of the line are ignored.</p>1027<p>This means that the two following regular expression objects that match a1028decimal number are functionally equal:</p>1029<div class="highlight-python3 notranslate"><div class="highlight"><pre><span></span><span class="n">a</span> <span class="o">=</span> <span class="n">re</span><span class="o">.</span><span class="n">compile</span><span class="p">(</span><span class="sa">r</span><span class="s2">"""\d + # the integral part</span>1030<span class="s2"> \. # the decimal point</span>1031<span class="s2"> \d * # some fractional digits"""</span><span class="p">,</span> <span class="n">re</span><span class="o">.</span><span class="n">X</span><span class="p">)</span>1032<span class="n">b</span> <span class="o">=</span> <span class="n">re</span><span class="o">.</span><span class="n">compile</span><span class="p">(</span><span class="sa">r</span><span class="s2">"\d+\.\d*"</span><span class="p">)</span>1033</pre></div>1034</div>1035<p>Corresponds to the inline flag <code class="docutils literal notranslate"><span class="pre">(?x)</span></code>.</p>1036</dd></dl>1037 1038</section>1039<section id="functions">1040<h3>Functions<a class="headerlink" href="#functions" title="Link to this heading">¶</a></h3>1041<dl class="py function">1042<dt class="sig sig-object py" id="re.compile">1043<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">compile</span></span><span class="sig-paren">(</span><em class="sig-param"><span class="n"><span class="pre">pattern</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">flags</span></span><span class="o"><span class="pre">=</span></span><span class="default_value"><span class="pre">0</span></span></em><span class="sig-paren">)</span><a class="headerlink" href="#re.compile" title="Link to this definition">¶</a></dt>1044<dd><p>Compile a regular expression pattern into a <a class="reference internal" href="#re-objects"><span class="std std-ref">regular expression object</span></a>, which can be used for matching using its1045<a class="reference internal" href="#re.Pattern.prefixmatch" title="re.Pattern.prefixmatch"><code class="xref py py-func docutils literal notranslate"><span class="pre">prefixmatch()</span></code></a>,1046<a class="reference internal" href="#re.Pattern.search" title="re.Pattern.search"><code class="xref py py-func docutils literal notranslate"><span class="pre">search()</span></code></a>, and other methods, described below.</p>1047<p>The expression’s behaviour can be modified by specifying a <em>flags</em> value.1048Values can be any of the <a class="reference internal" href="#flags">flags</a> variables, combined using bitwise OR1049(the <code class="docutils literal notranslate"><span class="pre">|</span></code> operator).</p>1050<p>The sequence</p>1051<div class="highlight-python3 notranslate"><div class="highlight"><pre><span></span><span class="n">prog</span> <span class="o">=</span> <span class="n">re</span><span class="o">.</span><span class="n">compile</span><span class="p">(</span><span class="n">pattern</span><span class="p">)</span>1052<span class="n">result</span> <span class="o">=</span> <span class="n">prog</span><span class="o">.</span><span class="n">search</span><span class="p">(</span><span class="n">string</span><span class="p">)</span>1053</pre></div>1054</div>1055<p>is equivalent to</p>1056<div class="highlight-python3 notranslate"><div class="highlight"><pre><span></span><span class="n">result</span> <span class="o">=</span> <span class="n">re</span><span class="o">.</span><span class="n">search</span><span class="p">(</span><span class="n">pattern</span><span class="p">,</span> <span class="n">string</span><span class="p">)</span>1057</pre></div>1058</div>1059<p>but using <a class="reference internal" href="#re.compile" title="re.compile"><code class="xref py py-func docutils literal notranslate"><span class="pre">re.compile()</span></code></a> and saving the resulting regular expression1060object for reuse is more efficient when the expression will be used several1061times in a single program.</p>1062<div class="admonition note">1063<p class="admonition-title">Note</p>1064<p>The compiled versions of the most recent patterns passed to1065<a class="reference internal" href="#re.compile" title="re.compile"><code class="xref py py-func docutils literal notranslate"><span class="pre">re.compile()</span></code></a> and the module-level matching functions are cached, so1066programs that use only a few regular expressions at a time needn’t worry1067about compiling regular expressions.</p>1068</div>1069</dd></dl>1070 1071<dl class="py function">1072<dt class="sig sig-object py" id="re.search">1073<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">search</span></span><span class="sig-paren">(</span><em class="sig-param"><span class="n"><span class="pre">pattern</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">string</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">flags</span></span><span class="o"><span class="pre">=</span></span><span class="default_value"><span class="pre">0</span></span></em><span class="sig-paren">)</span><a class="headerlink" href="#re.search" title="Link to this definition">¶</a></dt>1074<dd><p>Scan through <em>string</em> looking for the first location where the regular expression1075<em>pattern</em> produces a match, and return a corresponding <a class="reference internal" href="#re.Match" title="re.Match"><code class="xref py py-class docutils literal notranslate"><span class="pre">Match</span></code></a>. Return1076<code class="docutils literal notranslate"><span class="pre">None</span></code> if no position in the string matches the pattern; note that this is1077different from finding a zero-length match at some point in the string.</p>1078<p>The expression’s behaviour can be modified by specifying a <em>flags</em> value.1079Values can be any of the <a class="reference internal" href="#flags">flags</a> variables, combined using bitwise OR1080(the <code class="docutils literal notranslate"><span class="pre">|</span></code> operator).</p>1081</dd></dl>1082 1083<dl class="py function">1084<dt class="sig sig-object py" id="re.prefixmatch">1085<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">prefixmatch</span></span><span class="sig-paren">(</span><em class="sig-param"><span class="n"><span class="pre">pattern</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">string</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">flags</span></span><span class="o"><span class="pre">=</span></span><span class="default_value"><span class="pre">0</span></span></em><span class="sig-paren">)</span><a class="headerlink" href="#re.prefixmatch" title="Link to this definition">¶</a></dt>1086<dd></dd></dl>1087 1088<dl class="py function">1089<dt class="sig sig-object py" id="re.match">1090<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">match</span></span><span class="sig-paren">(</span><em class="sig-param"><span class="n"><span class="pre">pattern</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">string</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">flags</span></span><span class="o"><span class="pre">=</span></span><span class="default_value"><span class="pre">0</span></span></em><span class="sig-paren">)</span><a class="headerlink" href="#re.match" title="Link to this definition">¶</a></dt>1091<dd><p>If zero or more characters at the beginning of <em>string</em> match the regular1092expression <em>pattern</em>, return a corresponding <a class="reference internal" href="#re.Match" title="re.Match"><code class="xref py py-class docutils literal notranslate"><span class="pre">Match</span></code></a>. Return1093<code class="docutils literal notranslate"><span class="pre">None</span></code> if the string does not match the pattern; note that this is1094different from a zero-length match.</p>1095<div class="admonition note">1096<p class="admonition-title">Note</p>1097<p>Even in <a class="reference internal" href="#re.MULTILINE" title="re.MULTILINE"><code class="xref py py-const docutils literal notranslate"><span class="pre">MULTILINE</span></code></a> mode, this will only match at the1098beginning of the string and not at the beginning of each line.</p>1099</div>1100<p>If you want to locate a match anywhere in <em>string</em>, use <a class="reference internal" href="#re.search" title="re.search"><code class="xref py py-func docutils literal notranslate"><span class="pre">search()</span></code></a>1101instead (see also <a class="reference internal" href="#search-vs-match"><span class="std std-ref">search() vs. prefixmatch()</span></a>).</p>1102<p>The expression’s behaviour can be modified by specifying a <em>flags</em> value.1103Values can be any of the <a class="reference internal" href="#flags">flags</a> variables, combined using bitwise OR1104(the <code class="docutils literal notranslate"><span class="pre">|</span></code> operator).</p>1105<p>This function now has two names and has long been known as1106<a class="reference internal" href="#re.match" title="re.match"><code class="xref py py-func docutils literal notranslate"><span class="pre">match()</span></code></a>. Use that name when you need to retain compatibility with1107older Python versions.</p>1108<div class="versionchanged">1109<p><span class="versionmodified changed">Changed in version 3.15.0a6 (unreleased): </span>The alternate <a class="reference internal" href="#re.prefixmatch" title="re.prefixmatch"><code class="xref py py-func docutils literal notranslate"><span class="pre">prefixmatch()</span></code></a> name of this API was added as a1110more explicitly descriptive name than <a class="reference internal" href="#re.match" title="re.match"><code class="xref py py-func docutils literal notranslate"><span class="pre">match()</span></code></a>. Use it to better1111express intent. The norm in other languages and regular expression1112implementations is to use the term <em>match</em> to refer to the behavior of1113what Python has always called <a class="reference internal" href="#re.search" title="re.search"><code class="xref py py-func docutils literal notranslate"><span class="pre">search()</span></code></a>.1114See <a class="reference internal" href="#prefixmatch-vs-match"><span class="std std-ref">prefixmatch() vs. match()</span></a>.</p>1115</div>1116</dd></dl>1117 1118<dl class="py function">1119<dt class="sig sig-object py" id="re.fullmatch">1120<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">fullmatch</span></span><span class="sig-paren">(</span><em class="sig-param"><span class="n"><span class="pre">pattern</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">string</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">flags</span></span><span class="o"><span class="pre">=</span></span><span class="default_value"><span class="pre">0</span></span></em><span class="sig-paren">)</span><a class="headerlink" href="#re.fullmatch" title="Link to this definition">¶</a></dt>1121<dd><p>If the whole <em>string</em> matches the regular expression <em>pattern</em>, return a1122corresponding <a class="reference internal" href="#re.Match" title="re.Match"><code class="xref py py-class docutils literal notranslate"><span class="pre">Match</span></code></a>. Return <code class="docutils literal notranslate"><span class="pre">None</span></code> if the string does not match1123the pattern; note that this is different from a zero-length match.</p>1124<p>The expression’s behaviour can be modified by specifying a <em>flags</em> value.1125Values can be any of the <a class="reference internal" href="#flags">flags</a> variables, combined using bitwise OR1126(the <code class="docutils literal notranslate"><span class="pre">|</span></code> operator).</p>1127<div class="versionadded">1128<p><span class="versionmodified added">Added in version 3.4.</span></p>1129</div>1130</dd></dl>1131 1132<dl class="py function">1133<dt class="sig sig-object py" id="re.split">1134<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">split</span></span><span class="sig-paren">(</span><em class="sig-param"><span class="n"><span class="pre">pattern</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">string</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">maxsplit</span></span><span class="o"><span class="pre">=</span></span><span class="default_value"><span class="pre">0</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">flags</span></span><span class="o"><span class="pre">=</span></span><span class="default_value"><span class="pre">0</span></span></em><span class="sig-paren">)</span><a class="headerlink" href="#re.split" title="Link to this definition">¶</a></dt>1135<dd><p>Split <em>string</em> by the occurrences of <em>pattern</em>. If capturing parentheses are1136used in <em>pattern</em>, then the text of all groups in the pattern are also returned1137as part of the resulting list. If <em>maxsplit</em> is nonzero, at most <em>maxsplit</em>1138splits occur, and the remainder of the string is returned as the final element1139of the list.</p>1140<div class="highlight-python3 notranslate"><div class="highlight"><pre><span></span><span class="gp">>>> </span><span class="n">re</span><span class="o">.</span><span class="n">split</span><span class="p">(</span><span class="sa">r</span><span class="s1">'\W+'</span><span class="p">,</span> <span class="s1">'Words, words, words.'</span><span class="p">)</span>1141<span class="go">['Words', 'words', 'words', '']</span>1142<span class="gp">>>> </span><span class="n">re</span><span class="o">.</span><span class="n">split</span><span class="p">(</span><span class="sa">r</span><span class="s1">'(\W+)'</span><span class="p">,</span> <span class="s1">'Words, words, words.'</span><span class="p">)</span>1143<span class="go">['Words', ', ', 'words', ', ', 'words', '.', '']</span>1144<span class="gp">>>> </span><span class="n">re</span><span class="o">.</span><span class="n">split</span><span class="p">(</span><span class="sa">r</span><span class="s1">'\W+'</span><span class="p">,</span> <span class="s1">'Words, words, words.'</span><span class="p">,</span> <span class="n">maxsplit</span><span class="o">=</span><span class="mi">1</span><span class="p">)</span>1145<span class="go">['Words', 'words, words.']</span>1146<span class="gp">>>> </span><span class="n">re</span><span class="o">.</span><span class="n">split</span><span class="p">(</span><span class="s1">'[a-f]+'</span><span class="p">,</span> <span class="s1">'0a3B9'</span><span class="p">,</span> <span class="n">flags</span><span class="o">=</span><span class="n">re</span><span class="o">.</span><span class="n">IGNORECASE</span><span class="p">)</span>1147<span class="go">['0', '3', '9']</span>1148</pre></div>1149</div>1150<p>If there are capturing groups in the separator and it matches at the start of1151the string, the result will start with an empty string. The same holds for1152the end of the string:</p>1153<div class="highlight-python3 notranslate"><div class="highlight"><pre><span></span><span class="gp">>>> </span><span class="n">re</span><span class="o">.</span><span class="n">split</span><span class="p">(</span><span class="sa">r</span><span class="s1">'(\W+)'</span><span class="p">,</span> <span class="s1">'...words, words...'</span><span class="p">)</span>1154<span class="go">['', '...', 'words', ', ', 'words', '...', '']</span>1155</pre></div>1156</div>1157<p>That way, separator components are always found at the same relative1158indices within the result list.</p>1159<p>Adjacent empty matches are not possible, but an empty match can occur1160immediately after a non-empty match.</p>1161<div class="highlight-pycon notranslate"><div class="highlight"><pre><span></span><span class="gp">>>> </span><span class="n">re</span><span class="o">.</span><span class="n">split</span><span class="p">(</span><span class="sa">r</span><span class="s1">'\b'</span><span class="p">,</span> <span class="s1">'Words, words, words.'</span><span class="p">)</span>1162<span class="go">['', 'Words', ', ', 'words', ', ', 'words', '.']</span>1163<span class="gp">>>> </span><span class="n">re</span><span class="o">.</span><span class="n">split</span><span class="p">(</span><span class="sa">r</span><span class="s1">'\W*'</span><span class="p">,</span> <span class="s1">'...words...'</span><span class="p">)</span>1164<span class="go">['', '', 'w', 'o', 'r', 'd', 's', '', '']</span>1165<span class="gp">>>> </span><span class="n">re</span><span class="o">.</span><span class="n">split</span><span class="p">(</span><span class="sa">r</span><span class="s1">'(\W*)'</span><span class="p">,</span> <span class="s1">'...words...'</span><span class="p">)</span>1166<span class="go">['', '...', '', '', 'w', '', 'o', '', 'r', '', 'd', '', 's', '...', '', '', '']</span>1167</pre></div>1168</div>1169<p>The expression’s behaviour can be modified by specifying a <em>flags</em> value.1170Values can be any of the <a class="reference internal" href="#flags">flags</a> variables, combined using bitwise OR1171(the <code class="docutils literal notranslate"><span class="pre">|</span></code> operator).</p>1172<div class="versionchanged">1173<p><span class="versionmodified changed">Changed in version 3.1: </span>Added the optional flags argument.</p>1174</div>1175<div class="versionchanged">1176<p><span class="versionmodified changed">Changed in version 3.7: </span>Added support of splitting on a pattern that could match an empty string.</p>1177</div>1178<div class="deprecated">1179<p><span class="versionmodified deprecated">Deprecated since version 3.13: </span>Passing <em>maxsplit</em> and <em>flags</em> as positional arguments is deprecated.1180In future Python versions they will be1181<a class="reference internal" href="../glossary.html#keyword-only-parameter"><span class="std std-ref">keyword-only parameters</span></a>.</p>1182</div>1183</dd></dl>1184 1185<dl class="py function">1186<dt class="sig sig-object py" id="re.findall">1187<span class="sig-prename descclassname"><span class="pre">re.</span></span><span class="sig-name descname"><span class="pre">findall</span></span><span class="sig-paren">(</span><em class="sig-param"><span class="n"><span class="pre">pattern</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">string</span></span></em>, <em class="sig-param"><span class="n"><span class="pre">flags</span></span><span class="o"><span class="pre">=</span></span><span class="default_value"><span class="pre">0</span></span></em><span class="sig-paren">)</span><a class="headerlink" href="#re.findall" title="Link to this definition">¶</a></dt>1188<dd><p>Return all non-overlapping matches of <em>pattern</em> in <em>string</em>, as a list of1189strings or tuples. The <em>string</em> is scanned left-to-right, and matches1190are returned in the order found. Empty matches are included in the result.</p>1191<p>The result depends on the number of capturing groups in the pattern.1192If there are no groups, return a list of strings matching the whole1193pattern. If there is exactly one group, return a list of strings1194matching that group. If multiple groups are present, return a list1195of tuples of strings matching the groups. Non-capturing groups do not1196affect the form of the result.</p>1197<div class="doctest highlight-default notranslate"><div class="highlight"><pre><span></span><span class="gp">>>> </span><span class="n">re</span><span class="o">.</span><span class="n">findall</span><span class="p">(</span><span class="sa">r</span><span class="s1">'\bf[a-z]*'</span><span class="p">,</span> <span class="s1">'which foot or hand fell fastest'</span><span class="p">)</span>1198<span class="go">['foot', 'fell', 'fastest']</span>1199<span class="gp">>>> </span><span class="n">re</span><span class="o">.</span><span class="n">findall</span><span class="p">(</span><span class="sa">r</span><span class="s1">'(\w+)=(\d+)'</span><span class="p">,</span> <span class="s1">'set width=20 and height=10'</span><span class="p">)</span>1200<span class="go">[('width', '20'), ('height', '10')]</span>