Team Ai
Modelpublic

AryaWu/sqlite

sourceHugging Faceupdated 10mo agoView on Hugging Face
0likes
wal.c4622 linesDownload Raw Back to src
1/*2** 2010 February 13**4** The author disclaims copyright to this source code.  In place of5** a legal notice, here is a blessing:6**7**    May you do good and not evil.8**    May you find forgiveness for yourself and forgive others.9**    May you share freely, never taking more than you give.10**11*************************************************************************12**13** This file contains the implementation of a write-ahead log (WAL) used in14** "journal_mode=WAL" mode.15**16** WRITE-AHEAD LOG (WAL) FILE FORMAT17**18** A WAL file consists of a header followed by zero or more "frames".19** Each frame records the revised content of a single page from the20** database file.  All changes to the database are recorded by writing21** frames into the WAL.  Transactions commit when a frame is written that22** contains a commit marker.  A single WAL can and usually does record23** multiple transactions.  Periodically, the content of the WAL is24** transferred back into the database file in an operation called a25** "checkpoint".26**27** A single WAL file can be used multiple times.  In other words, the28** WAL can fill up with frames and then be checkpointed and then new29** frames can overwrite the old ones.  A WAL always grows from beginning30** toward the end.  Checksums and counters attached to each frame are31** used to determine which frames within the WAL are valid and which32** are leftovers from prior checkpoints.33**34** The WAL header is 32 bytes in size and consists of the following eight35** big-endian 32-bit unsigned integer values:36**37**     0: Magic number.  0x377f0682 or 0x377f068338**     4: File format version.  Currently 300700039**     8: Database page size.  Example: 102440**    12: Checkpoint sequence number41**    16: Salt-1, random integer incremented with each checkpoint42**    20: Salt-2, a different random integer changing with each ckpt43**    24: Checksum-1 (first part of checksum for first 24 bytes of header).44**    28: Checksum-2 (second part of checksum for first 24 bytes of header).45**46** Immediately following the wal-header are zero or more frames. Each47** frame consists of a 24-byte frame-header followed by <page-size> bytes48** of page data. The frame-header is six big-endian 32-bit unsigned49** integer values, as follows:50**51**     0: Page number.52**     4: For commit records, the size of the database image in pages53**        after the commit. For all other records, zero.54**     8: Salt-1 (copied from the header)55**    12: Salt-2 (copied from the header)56**    16: Checksum-1.57**    20: Checksum-2.58**59** A frame is considered valid if and only if the following conditions are60** true:61**62**    (1) The salt-1 and salt-2 values in the frame-header match63**        salt values in the wal-header64**65**    (2) The checksum values in the final 8 bytes of the frame-header66**        exactly match the checksum computed consecutively on the67**        WAL header and the first 8 bytes and the content of all frames68**        up to and including the current frame.69**70** The checksum is computed using 32-bit big-endian integers if the71** magic number in the first 4 bytes of the WAL is 0x377f0683 and it72** is computed using little-endian if the magic number is 0x377f0682.73** The checksum values are always stored in the frame header in a74** big-endian format regardless of which byte order is used to compute75** the checksum.  The checksum is computed by interpreting the input as76** an even number of unsigned 32-bit integers: x[0] through x[N].  The77** algorithm used for the checksum is as follows:78**79**   for i from 0 to n-1 step 2:80**     s0 += x[i] + s1;81**     s1 += x[i+1] + s0;82**   endfor83**84** Note that s0 and s1 are both weighted checksums using fibonacci weights85** in reverse order (the largest fibonacci weight occurs on the first element86** of the sequence being summed.)  The s1 value spans all 32-bit87** terms of the sequence whereas s0 omits the final term.88**89** On a checkpoint, the WAL is first VFS.xSync-ed, then valid content of the90** WAL is transferred into the database, then the database is VFS.xSync-ed.91** The VFS.xSync operations serve as write barriers - all writes launched92** before the xSync must complete before any write that launches after the93** xSync begins.94**95** After each checkpoint, the salt-1 value is incremented and the salt-296** value is randomized.  This prevents old and new frames in the WAL from97** being considered valid at the same time and being checkpointing together98** following a crash.99**100** READER ALGORITHM101**102** To read a page from the database (call it page number P), a reader103** first checks the WAL to see if it contains page P.  If so, then the104** last valid instance of page P that is a followed by a commit frame105** or is a commit frame itself becomes the value read.  If the WAL106** contains no copies of page P that are valid and which are a commit107** frame or are followed by a commit frame, then page P is read from108** the database file.109**110** To start a read transaction, the reader records the index of the last111** valid frame in the WAL.  The reader uses this recorded "mxFrame" value112** for all subsequent read operations.  New transactions can be appended113** to the WAL, but as long as the reader uses its original mxFrame value114** and ignores the newly appended content, it will see a consistent snapshot115** of the database from a single point in time.  This technique allows116** multiple concurrent readers to view different versions of the database117** content simultaneously.118**119** The reader algorithm in the previous paragraphs works correctly, but120** because frames for page P can appear anywhere within the WAL, the121** reader has to scan the entire WAL looking for page P frames.  If the122** WAL is large (multiple megabytes is typical) that scan can be slow,123** and read performance suffers.  To overcome this problem, a separate124** data structure called the wal-index is maintained to expedite the125** search for frames of a particular page.126**127** WAL-INDEX FORMAT128**129** Conceptually, the wal-index is shared memory, though VFS implementations130** might choose to implement the wal-index using a mmapped file.  Because131** the wal-index is shared memory, SQLite does not support journal_mode=WAL132** on a network filesystem.  All users of the database must be able to133** share memory.134**135** In the default unix and windows implementation, the wal-index is a mmapped136** file whose name is the database name with a "-shm" suffix added.  For that137** reason, the wal-index is sometimes called the "shm" file.138**139** The wal-index is transient.  After a crash, the wal-index can (and should140** be) reconstructed from the original WAL file.  In fact, the VFS is required141** to either truncate or zero the header of the wal-index when the last142** connection to it closes.  Because the wal-index is transient, it can143** use an architecture-specific format; it does not have to be cross-platform.144** Hence, unlike the database and WAL file formats which store all values145** as big endian, the wal-index can store multi-byte values in the native146** byte order of the host computer.147**148** The purpose of the wal-index is to answer this question quickly:  Given149** a page number P and a maximum frame index M, return the index of the150** last frame in the wal before frame M for page P in the WAL, or return151** NULL if there are no frames for page P in the WAL prior to M.152**153** The wal-index consists of a header region, followed by an one or154** more index blocks.155**156** The wal-index header contains the total number of frames within the WAL157** in the mxFrame field.158**159** Each index block except for the first contains information on160** HASHTABLE_NPAGE frames. The first index block contains information on161** HASHTABLE_NPAGE_ONE frames. The values of HASHTABLE_NPAGE_ONE and162** HASHTABLE_NPAGE are selected so that together the wal-index header and163** first index block are the same size as all other index blocks in the164** wal-index.  The values are:165**166**   HASHTABLE_NPAGE      4096167**   HASHTABLE_NPAGE_ONE  4062168**169** Each index block contains two sections, a page-mapping that contains the170** database page number associated with each wal frame, and a hash-table171** that allows readers to query an index block for a specific page number.172** The page-mapping is an array of HASHTABLE_NPAGE (or HASHTABLE_NPAGE_ONE173** for the first index block) 32-bit page numbers. The first entry in the174** first index-block contains the database page number corresponding to the175** first frame in the WAL file. The first entry in the second index block176** in the WAL file corresponds to the (HASHTABLE_NPAGE_ONE+1)th frame in177** the log, and so on.178**179** The last index block in a wal-index usually contains less than the full180** complement of HASHTABLE_NPAGE (or HASHTABLE_NPAGE_ONE) page-numbers,181** depending on the contents of the WAL file. This does not change the182** allocated size of the page-mapping array - the page-mapping array merely183** contains unused entries.184**185** Even without using the hash table, the last frame for page P186** can be found by scanning the page-mapping sections of each index block187** starting with the last index block and moving toward the first, and188** within each index block, starting at the end and moving toward the189** beginning.  The first entry that equals P corresponds to the frame190** holding the content for that page.191**192** The hash table consists of HASHTABLE_NSLOT 16-bit unsigned integers.193** HASHTABLE_NSLOT = 2*HASHTABLE_NPAGE, and there is one entry in the194** hash table for each page number in the mapping section, so the hash195** table is never more than half full.  The expected number of collisions196** prior to finding a match is 1.  Each entry of the hash table is an197** 1-based index of an entry in the mapping section of the same198** index block.   Let K be the 1-based index of the largest entry in199** the mapping section.  (For index blocks other than the last, K will200** always be exactly HASHTABLE_NPAGE (4096) and for the last index block201** K will be (mxFrame%HASHTABLE_NPAGE).)  Unused slots of the hash table202** contain a value of 0.203**204** To look for page P in the hash table, first compute a hash iKey on205** P as follows:206**207**      iKey = (P * 383) % HASHTABLE_NSLOT208**209** Then start scanning entries of the hash table, starting with iKey210** (wrapping around to the beginning when the end of the hash table is211** reached) until an unused hash slot is found. Let the first unused slot212** be at index iUnused.  (iUnused might be less than iKey if there was213** wrap-around.) Because the hash table is never more than half full,214** the search is guaranteed to eventually hit an unused entry.  Let215** iMax be the value between iKey and iUnused, closest to iUnused,216** where aHash[iMax]==P.  If there is no iMax entry (if there exists217** no hash slot such that aHash[i]==p) then page P is not in the218** current index block.  Otherwise the iMax-th mapping entry of the219** current index block corresponds to the last entry that references220** page P.221**222** A hash search begins with the last index block and moves toward the223** first index block, looking for entries corresponding to page P.  On224** average, only two or three slots in each index block need to be225** examined in order to either find the last entry for page P, or to226** establish that no such entry exists in the block.  Each index block227** holds over 4000 entries.  So two or three index blocks are sufficient228** to cover a typical 10 megabyte WAL file, assuming 1K pages.  8 or 10229** comparisons (on average) suffice to either locate a frame in the230** WAL or to establish that the frame does not exist in the WAL.  This231** is much faster than scanning the entire 10MB WAL.232**233** Note that entries are added in order of increasing K.  Hence, one234** reader might be using some value K0 and a second reader that started235** at a later time (after additional transactions were added to the WAL236** and to the wal-index) might be using a different value K1, where K1>K0.237** Both readers can use the same hash table and mapping section to get238** the correct result.  There may be entries in the hash table with239** K>K0 but to the first reader, those entries will appear to be unused240** slots in the hash table and so the first reader will get an answer as241** if no values greater than K0 had ever been inserted into the hash table242** in the first place - which is what reader one wants.  Meanwhile, the243** second reader using K1 will see additional values that were inserted244** later, which is exactly what reader two wants.245**246** When a rollback occurs, the value of K is decreased. Hash table entries247** that correspond to frames greater than the new K value are removed248** from the hash table at this point.249*/250#ifndef SQLITE_OMIT_WAL251 252#include "wal.h"253 254/*255** Trace output macros256*/257#if defined(SQLITE_TEST) && defined(SQLITE_DEBUG)258int sqlite3WalTrace = 0;259# define WALTRACE(X)  if(sqlite3WalTrace) sqlite3DebugPrintf X260#else261# define WALTRACE(X)262#endif263 264/*265** The maximum (and only) versions of the wal and wal-index formats266** that may be interpreted by this version of SQLite.267**268** If a client begins recovering a WAL file and finds that (a) the checksum269** values in the wal-header are correct and (b) the version field is not270** WAL_MAX_VERSION, recovery fails and SQLite returns SQLITE_CANTOPEN.271**272** Similarly, if a client successfully reads a wal-index header (i.e. the273** checksum test is successful) and finds that the version field is not274** WALINDEX_MAX_VERSION, then no read-transaction is opened and SQLite275** returns SQLITE_CANTOPEN.276*/277#define WAL_MAX_VERSION      3007000278#define WALINDEX_MAX_VERSION 3007000279 280/*281** Index numbers for various locking bytes.   WAL_NREADER is the number282** of available reader locks and should be at least 3.  The default283** is SQLITE_SHM_NLOCK==8 and  WAL_NREADER==5.284**285** Technically, the various VFSes are free to implement these locks however286** they see fit.  However, compatibility is encouraged so that VFSes can287** interoperate.  The standard implementation used on both unix and windows288** is for the index number to indicate a byte offset into the289** WalCkptInfo.aLock[] array in the wal-index header.  In other words, all290** locks are on the shm file.  The WALINDEX_LOCK_OFFSET constant (which291** should be 120) is the location in the shm file for the first locking292** byte.293*/294#define WAL_WRITE_LOCK         0295#define WAL_ALL_BUT_WRITE      1296#define WAL_CKPT_LOCK          1297#define WAL_RECOVER_LOCK       2298#define WAL_READ_LOCK(I)       (3+(I))299#define WAL_NREADER            (SQLITE_SHM_NLOCK-3)300 301 302/* Object declarations */303typedef struct WalIndexHdr WalIndexHdr;304typedef struct WalIterator WalIterator;305typedef struct WalCkptInfo WalCkptInfo;306 307 308/*309** The following object holds a copy of the wal-index header content.310**311** The actual header in the wal-index consists of two copies of this312** object followed by one instance of the WalCkptInfo object.313** For all versions of SQLite through 3.10.0 and probably beyond,314** the locking bytes (WalCkptInfo.aLock) start at offset 120 and315** the total header size is 136 bytes.316**317** The szPage value can be any power of 2 between 512 and 32768, inclusive.318** Or it can be 1 to represent a 65536-byte page.  The latter case was319** added in 3.7.1 when support for 64K pages was added.320*/321struct WalIndexHdr {322  u32 iVersion;                   /* Wal-index version */323  u32 unused;                     /* Unused (padding) field */324  u32 iChange;                    /* Counter incremented each transaction */325  u8 isInit;                      /* 1 when initialized */326  u8 bigEndCksum;                 /* True if checksums in WAL are big-endian */327  u16 szPage;                     /* Database page size in bytes. 1==64K */328  u32 mxFrame;                    /* Index of last valid frame in the WAL */329  u32 nPage;                      /* Size of database in pages */330  u32 aFrameCksum[2];             /* Checksum of last frame in log */331  u32 aSalt[2];                   /* Two salt values copied from WAL header */332  u32 aCksum[2];                  /* Checksum over all prior fields */333};334 335/*336** A copy of the following object occurs in the wal-index immediately337** following the second copy of the WalIndexHdr.  This object stores338** information used by checkpoint.339**340** nBackfill is the number of frames in the WAL that have been written341** back into the database. (We call the act of moving content from WAL to342** database "backfilling".)  The nBackfill number is never greater than343** WalIndexHdr.mxFrame.  nBackfill can only be increased by threads344** holding the WAL_CKPT_LOCK lock (which includes a recovery thread).345** However, a WAL_WRITE_LOCK thread can move the value of nBackfill from346** mxFrame back to zero when the WAL is reset.347**348** nBackfillAttempted is the largest value of nBackfill that a checkpoint349** has attempted to achieve.  Normally nBackfill==nBackfillAtempted, however350** the nBackfillAttempted is set before any backfilling is done and the351** nBackfill is only set after all backfilling completes.  So if a checkpoint352** crashes, nBackfillAttempted might be larger than nBackfill.  The353** WalIndexHdr.mxFrame must never be less than nBackfillAttempted.354**355** The aLock[] field is a set of bytes used for locking.  These bytes should356** never be read or written.357**358** There is one entry in aReadMark[] for each reader lock.  If a reader359** holds read-lock K, then the value in aReadMark[K] is no greater than360** the mxFrame for that reader.  The value READMARK_NOT_USED (0xffffffff)361** for any aReadMark[] means that entry is unused.  aReadMark[0] is362** a special case; its value is never used and it exists as a place-holder363** to avoid having to offset aReadMark[] indexes by one.  Readers holding364** WAL_READ_LOCK(0) always ignore the entire WAL and read all content365** directly from the database.366**367** The value of aReadMark[K] may only be changed by a thread that368** is holding an exclusive lock on WAL_READ_LOCK(K).  Thus, the value of369** aReadMark[K] cannot changed while there is a reader is using that mark370** since the reader will be holding a shared lock on WAL_READ_LOCK(K).371**372** The checkpointer may only transfer frames from WAL to database where373** the frame numbers are less than or equal to every aReadMark[] that is374** in use (that is, every aReadMark[j] for which there is a corresponding375** WAL_READ_LOCK(j)).  New readers (usually) pick the aReadMark[] with the376** largest value and will increase an unused aReadMark[] to mxFrame if there377** is not already an aReadMark[] equal to mxFrame.  The exception to the378** previous sentence is when nBackfill equals mxFrame (meaning that everything379** in the WAL has been backfilled into the database) then new readers380** will choose aReadMark[0] which has value 0 and hence such reader will381** get all their all content directly from the database file and ignore382** the WAL.383**384** Writers normally append new frames to the end of the WAL.  However,385** if nBackfill equals mxFrame (meaning that all WAL content has been386** written back into the database) and if no readers are using the WAL387** (in other words, if there are no WAL_READ_LOCK(i) where i>0) then388** the writer will first "reset" the WAL back to the beginning and start389** writing new content beginning at frame 1.390**391** We assume that 32-bit loads are atomic and so no locks are needed in392** order to read from any aReadMark[] entries.393*/394struct WalCkptInfo {395  u32 nBackfill;                  /* Number of WAL frames backfilled into DB */396  u32 aReadMark[WAL_NREADER];     /* Reader marks */397  u8 aLock[SQLITE_SHM_NLOCK];     /* Reserved space for locks */398  u32 nBackfillAttempted;         /* WAL frames perhaps written, or maybe not */399  u32 notUsed0;                   /* Available for future enhancements */400};401#define READMARK_NOT_USED  0xffffffff402 403/*404** This is a schematic view of the complete 136-byte header of the405** wal-index file (also known as the -shm file):406**407**      +-----------------------------+408**   0: | iVersion                    | \409**      +-----------------------------+  |410**   4: | (unused padding)            |  |411**      +-----------------------------+  |412**   8: | iChange                     |  |413**      +-------+-------+-------------+  |414**  12: | bInit |  bBig |   szPage    |  |415**      +-------+-------+-------------+  |416**  16: | mxFrame                     |  |  First copy of the417**      +-----------------------------+  |  WalIndexHdr object418**  20: | nPage                       |  |419**      +-----------------------------+  |420**  24: | aFrameCksum                 |  |421**      |                             |  |422**      +-----------------------------+  |423**  32: | aSalt                       |  |424**      |                             |  |425**      +-----------------------------+  |426**  40: | aCksum                      |  |427**      |                             | /428**      +-----------------------------+429**  48: | iVersion                    | \430**      +-----------------------------+  |431**  52: | (unused padding)            |  |432**      +-----------------------------+  |433**  56: | iChange                     |  |434**      +-------+-------+-------------+  |435**  60: | bInit |  bBig |   szPage    |  |436**      +-------+-------+-------------+  |  Second copy of the437**  64: | mxFrame                     |  |  WalIndexHdr438**      +-----------------------------+  |439**  68: | nPage                       |  |440**      +-----------------------------+  |441**  72: | aFrameCksum                 |  |442**      |                             |  |443**      +-----------------------------+  |444**  80: | aSalt                       |  |445**      |                             |  |446**      +-----------------------------+  |447**  88: | aCksum                      |  |448**      |                             | /449**      +-----------------------------+450**  96: | nBackfill                   |451**      +-----------------------------+452** 100: | 5 read marks                |453**      |                             |454**      |                             |455**      |                             |456**      |                             |457**      +-------+-------+------+------+458** 120: | Write | Ckpt  | Rcvr | Rd0  | \459**      +-------+-------+------+------+  ) 8 lock bytes460**      | Read1 | Read2 | Rd3  | Rd4  | /461**      +-------+-------+------+------+462** 128: | nBackfillAttempted          |463**      +-----------------------------+464** 132: | (unused padding)            |465**      +-----------------------------+466*/467 468/* A block of WALINDEX_LOCK_RESERVED bytes beginning at469** WALINDEX_LOCK_OFFSET is reserved for locks. Since some systems470** only support mandatory file-locks, we do not read or write data471** from the region of the file on which locks are applied.472*/473#define WALINDEX_LOCK_OFFSET (sizeof(WalIndexHdr)*2+offsetof(WalCkptInfo,aLock))474#define WALINDEX_HDR_SIZE    (sizeof(WalIndexHdr)*2+sizeof(WalCkptInfo))475 476/* Size of header before each frame in wal */477#define WAL_FRAME_HDRSIZE 24478 479/* Size of write ahead log header, including checksum. */480#define WAL_HDRSIZE 32481 482/* WAL magic value. Either this value, or the same value with the least483** significant bit also set (WAL_MAGIC | 0x00000001) is stored in 32-bit484** big-endian format in the first 4 bytes of a WAL file.485**486** If the LSB is set, then the checksums for each frame within the WAL487** file are calculated by treating all data as an array of 32-bit488** big-endian words. Otherwise, they are calculated by interpreting489** all data as 32-bit little-endian words.490*/491#define WAL_MAGIC 0x377f0682492 493/*494** Return the offset of frame iFrame in the write-ahead log file,495** assuming a database page size of szPage bytes. The offset returned496** is to the start of the write-ahead log frame-header.497*/498#define walFrameOffset(iFrame, szPage) (                               \499  WAL_HDRSIZE + ((iFrame)-1)*(i64)((szPage)+WAL_FRAME_HDRSIZE)         \500)501 502/*503** An open write-ahead log file is represented by an instance of the504** following object.505**506** writeLock:507**   This is usually set to 1 whenever the WRITER lock is held. However,508**   if it is set to 2, then the WRITER lock is held but must be released509**   by walHandleException() if a SEH exception is thrown.510*/511struct Wal {512  sqlite3_vfs *pVfs;         /* The VFS used to create pDbFd */513  sqlite3_file *pDbFd;       /* File handle for the database file */514  sqlite3_file *pWalFd;      /* File handle for WAL file */515  u32 iCallback;             /* Value to pass to log callback (or 0) */516  i64 mxWalSize;             /* Truncate WAL to this size upon reset */517  int nWiData;               /* Size of array apWiData */518  int szFirstBlock;          /* Size of first block written to WAL file */519  volatile u32 **apWiData;   /* Pointer to wal-index content in memory */520  u32 szPage;                /* Database page size */521  i16 readLock;              /* Which read lock is being held.  -1 for none */522  u8 syncFlags;              /* Flags to use to sync header writes */523  u8 exclusiveMode;          /* Non-zero if connection is in exclusive mode */524  u8 writeLock;              /* True if in a write transaction */525  u8 ckptLock;               /* True if holding a checkpoint lock */526  u8 readOnly;               /* WAL_RDWR, WAL_RDONLY, or WAL_SHM_RDONLY */527  u8 truncateOnCommit;       /* True to truncate WAL file on commit */528  u8 syncHeader;             /* Fsync the WAL header if true */529  u8 padToSectorBoundary;    /* Pad transactions out to the next sector */530  u8 bShmUnreliable;         /* SHM content is read-only and unreliable */531  WalIndexHdr hdr;           /* Wal-index header for current transaction */532  u32 minFrame;              /* Ignore wal frames before this one */533  u32 iReCksum;              /* On commit, recalculate checksums from here */534  const char *zWalName;      /* Name of WAL file */535  u32 nCkpt;                 /* Checkpoint sequence counter in the wal-header */536#ifdef SQLITE_USE_SEH537  u32 lockMask;              /* Mask of locks held */538  void *pFree;               /* Pointer to sqlite3_free() if exception thrown */539  u32 *pWiValue;             /* Value to write into apWiData[iWiPg] */540  int iWiPg;                 /* Write pWiValue into apWiData[iWiPg] */541  int iSysErrno;             /* System error code following exception */542#endif543#ifdef SQLITE_DEBUG544  int nSehTry;               /* Number of nested SEH_TRY{} blocks */545  u8 lockError;              /* True if a locking error has occurred */546#endif547#ifdef SQLITE_ENABLE_SNAPSHOT548  WalIndexHdr *pSnapshot;    /* Start transaction here if not NULL */549  int bGetSnapshot;          /* Transaction opened for sqlite3_get_snapshot() */550#endif551#ifdef SQLITE_ENABLE_SETLK_TIMEOUT552  sqlite3 *db;553#endif554};555 556/*557** Candidate values for Wal.exclusiveMode.558*/559#define WAL_NORMAL_MODE     0560#define WAL_EXCLUSIVE_MODE  1561#define WAL_HEAPMEMORY_MODE 2562 563/*564** Possible values for WAL.readOnly565*/566#define WAL_RDWR        0    /* Normal read/write connection */567#define WAL_RDONLY      1    /* The WAL file is readonly */568#define WAL_SHM_RDONLY  2    /* The SHM file is readonly */569 570/*571** Each page of the wal-index mapping contains a hash-table made up of572** an array of HASHTABLE_NSLOT elements of the following type.573*/574typedef u16 ht_slot;575 576/*577** This structure is used to implement an iterator that loops through578** all frames in the WAL in database page order. Where two or more frames579** correspond to the same database page, the iterator visits only the580** frame most recently written to the WAL (in other words, the frame with581** the largest index).582**583** The internals of this structure are only accessed by:584**585**   walIteratorInit() - Create a new iterator,586**   walIteratorNext() - Step an iterator,587**   walIteratorFree() - Free an iterator.588**589** This functionality is used by the checkpoint code (see walCheckpoint()).590*/591struct WalIterator {592  u32 iPrior;                     /* Last result returned from the iterator */593  int nSegment;                   /* Number of entries in aSegment[] */594  struct WalSegment {595    int iNext;                    /* Next slot in aIndex[] not yet returned */596    ht_slot *aIndex;              /* i0, i1, i2... such that aPgno[iN] ascend */597    u32 *aPgno;                   /* Array of page numbers. */598    int nEntry;                   /* Nr. of entries in aPgno[] and aIndex[] */599    int iZero;                    /* Frame number associated with aPgno[0] */600  } aSegment[FLEXARRAY];          /* One for every 32KB page in the wal-index */601};602 603/* Size (in bytes) of a WalIterator object suitable for N or fewer segments */604#define SZ_WALITERATOR(N)  \605     (offsetof(WalIterator,aSegment)+(N)*sizeof(struct WalSegment))606 607/*608** Define the parameters of the hash tables in the wal-index file. There609** is a hash-table following every HASHTABLE_NPAGE page numbers in the610** wal-index.611**612** Changing any of these constants will alter the wal-index format and613** create incompatibilities.614*/615#define HASHTABLE_NPAGE      4096                 /* Must be power of 2 */616#define HASHTABLE_HASH_1     383                  /* Should be prime */617#define HASHTABLE_NSLOT      (HASHTABLE_NPAGE*2)  /* Must be a power of 2 */618 619/*620** The block of page numbers associated with the first hash-table in a621** wal-index is smaller than usual. This is so that there is a complete622** hash-table on each aligned 32KB page of the wal-index.623*/624#define HASHTABLE_NPAGE_ONE  (HASHTABLE_NPAGE - (WALINDEX_HDR_SIZE/sizeof(u32)))625 626/* The wal-index is divided into pages of WALINDEX_PGSZ bytes each. */627#define WALINDEX_PGSZ   (                                         \628    sizeof(ht_slot)*HASHTABLE_NSLOT + HASHTABLE_NPAGE*sizeof(u32) \629)630 631/*632** Structured Exception Handling (SEH) is a Windows-specific technique633** for catching exceptions raised while accessing memory-mapped files.634**635** The -DSQLITE_USE_SEH compile-time option means to use SEH to catch and636** deal with system-level errors that arise during WAL -shm file processing.637** Without this compile-time option, any system-level faults that appear638** while accessing the memory-mapped -shm file will cause a process-wide639** signal to be deliver, which will more than likely cause the entire640** process to exit.641*/642#ifdef SQLITE_USE_SEH643#include <Windows.h>644 645/* Beginning of a block of code in which an exception might occur */646# define SEH_TRY    __try { \647   assert( walAssertLockmask(pWal) && pWal->nSehTry==0 ); \648   VVA_ONLY(pWal->nSehTry++);649 650/* The end of a block of code in which an exception might occur */651# define SEH_EXCEPT(X) \652   VVA_ONLY(pWal->nSehTry--); \653   assert( pWal->nSehTry==0 ); \654   } __except( sehExceptionFilter(pWal, GetExceptionCode(), GetExceptionInformation() ) ){ X }655 656/* Simulate a memory-mapping fault in the -shm file for testing purposes */657# define SEH_INJECT_FAULT sehInjectFault(pWal) 658 659/*660** The second argument is the return value of GetExceptionCode() for the 661** current exception. Return EXCEPTION_EXECUTE_HANDLER if the exception code662** indicates that the exception may have been caused by accessing the *-shm 663** file mapping. Or EXCEPTION_CONTINUE_SEARCH otherwise.664*/665static int sehExceptionFilter(Wal *pWal, int eCode, EXCEPTION_POINTERS *p){666  VVA_ONLY(pWal->nSehTry--);667  if( eCode==EXCEPTION_IN_PAGE_ERROR ){668    if( p && p->ExceptionRecord && p->ExceptionRecord->NumberParameters>=3 ){669      /* From MSDN: For this type of exception, the first element of the670      ** ExceptionInformation[] array is a read-write flag - 0 if the exception671      ** was thrown while reading, 1 if while writing. The second element is672      ** the virtual address being accessed. The "third array element specifies673      ** the underlying NTSTATUS code that resulted in the exception". */674      pWal->iSysErrno = (int)p->ExceptionRecord->ExceptionInformation[2];675    }676    return EXCEPTION_EXECUTE_HANDLER;677  }678  return EXCEPTION_CONTINUE_SEARCH;679}680 681/*682** If one is configured, invoke the xTestCallback callback with 650 as683** the argument. If it returns true, throw the same exception that is684** thrown by the system if the *-shm file mapping is accessed after it685** has been invalidated.686*/687static void sehInjectFault(Wal *pWal){688  int res;689  assert( pWal->nSehTry>0 );690 691  res = sqlite3FaultSim(650);692  if( res!=0 ){693    ULONG_PTR aArg[3];694    aArg[0] = 0;695    aArg[1] = 0;696    aArg[2] = (ULONG_PTR)res;697    RaiseException(EXCEPTION_IN_PAGE_ERROR, 0, 3, (const ULONG_PTR*)aArg);698  }699}700 701/*702** There are two ways to use this macro. To set a pointer to be freed703** if an exception is thrown:704**705**   SEH_FREE_ON_ERROR(0, pPtr);706**707** and to cancel the same:708**709**   SEH_FREE_ON_ERROR(pPtr, 0);710**711** In the first case, there must not already be a pointer registered to712** be freed. In the second case, pPtr must be the registered pointer.713*/714#define SEH_FREE_ON_ERROR(X,Y) \715  assert( (X==0 || Y==0) && pWal->pFree==X ); pWal->pFree = Y716 717/*718** There are two ways to use this macro. To arrange for pWal->apWiData[iPg]719** to be set to pValue if an exception is thrown:720**721**   SEH_SET_ON_ERROR(iPg, pValue);722**723** and to cancel the same:724**725**   SEH_SET_ON_ERROR(0, 0);726*/727#define SEH_SET_ON_ERROR(X,Y)  pWal->iWiPg = X; pWal->pWiValue = Y728 729#else730# define SEH_TRY          VVA_ONLY(pWal->nSehTry++);731# define SEH_EXCEPT(X)    VVA_ONLY(pWal->nSehTry--); assert( pWal->nSehTry==0 );732# define SEH_INJECT_FAULT assert( pWal->nSehTry>0 );733# define SEH_FREE_ON_ERROR(X,Y)734# define SEH_SET_ON_ERROR(X,Y)735#endif /* ifdef SQLITE_USE_SEH */736 737 738/*739** Obtain a pointer to the iPage'th page of the wal-index. The wal-index740** is broken into pages of WALINDEX_PGSZ bytes. Wal-index pages are741** numbered from zero.742**743** If the wal-index is currently smaller the iPage pages then the size744** of the wal-index might be increased, but only if it is safe to do745** so.  It is safe to enlarge the wal-index if pWal->writeLock is true746** or pWal->exclusiveMode==WAL_HEAPMEMORY_MODE.747**748** Three possible result scenarios:749**750**   (1)  rc==SQLITE_OK    and *ppPage==Requested-Wal-Index-Page751**   (2)  rc>=SQLITE_ERROR and *ppPage==NULL752**   (3)  rc==SQLITE_OK    and *ppPage==NULL  // only if iPage==0753**754** Scenario (3) can only occur when pWal->writeLock is false and iPage==0755*/756static SQLITE_NOINLINE int walIndexPageRealloc(757  Wal *pWal,               /* The WAL context */758  int iPage,               /* The page we seek */759  volatile u32 **ppPage    /* Write the page pointer here */760){761  int rc = SQLITE_OK;762 763  /* Enlarge the pWal->apWiData[] array if required */764  if( pWal->nWiData<=iPage ){765    sqlite3_int64 nByte = sizeof(u32*)*(1+(i64)iPage);766    volatile u32 **apNew;767    apNew = (volatile u32 **)sqlite3Realloc((void *)pWal->apWiData, nByte);768    if( !apNew ){769      *ppPage = 0;770      return SQLITE_NOMEM_BKPT;771    }772    memset((void*)&apNew[pWal->nWiData], 0,773           sizeof(u32*)*(iPage+1-pWal->nWiData));774    pWal->apWiData = apNew;775    pWal->nWiData = iPage+1;776  }777 778  /* Request a pointer to the required page from the VFS */779  assert( pWal->apWiData[iPage]==0 );780  if( pWal->exclusiveMode==WAL_HEAPMEMORY_MODE ){781    pWal->apWiData[iPage] = (u32 volatile *)sqlite3MallocZero(WALINDEX_PGSZ);782    if( !pWal->apWiData[iPage] ) rc = SQLITE_NOMEM_BKPT;783  }else{784    rc = sqlite3OsShmMap(pWal->pDbFd, iPage, WALINDEX_PGSZ,785        pWal->writeLock, (void volatile **)&pWal->apWiData[iPage]786    );787    assert( pWal->apWiData[iPage]!=0788         || rc!=SQLITE_OK789         || (pWal->writeLock==0 && iPage==0) );790    testcase( pWal->apWiData[iPage]==0 && rc==SQLITE_OK );791    if( rc==SQLITE_OK ){792      if( iPage>0 && sqlite3FaultSim(600) ) rc = SQLITE_NOMEM;793    }else if( (rc&0xff)==SQLITE_READONLY ){794      pWal->readOnly |= WAL_SHM_RDONLY;795      if( rc==SQLITE_READONLY ){796        rc = SQLITE_OK;797      }798    }799  }800 801  *ppPage = pWal->apWiData[iPage];802  assert( iPage==0 || *ppPage || rc!=SQLITE_OK );803  return rc;804}805static int walIndexPage(806  Wal *pWal,               /* The WAL context */807  int iPage,               /* The page we seek */808  volatile u32 **ppPage    /* Write the page pointer here */809){810  SEH_INJECT_FAULT;811  if( pWal->nWiData<=iPage || (*ppPage = pWal->apWiData[iPage])==0 ){812    return walIndexPageRealloc(pWal, iPage, ppPage);813  }814  return SQLITE_OK;815}816 817/*818** Return a pointer to the WalCkptInfo structure in the wal-index.819*/820static volatile WalCkptInfo *walCkptInfo(Wal *pWal){821  assert( pWal->nWiData>0 && pWal->apWiData[0] );822  SEH_INJECT_FAULT;823  return (volatile WalCkptInfo*)&(pWal->apWiData[0][sizeof(WalIndexHdr)/2]);824}825 826/*827** Return a pointer to the WalIndexHdr structure in the wal-index.828*/829static volatile WalIndexHdr *walIndexHdr(Wal *pWal){830  assert( pWal->nWiData>0 && pWal->apWiData[0] );831  SEH_INJECT_FAULT;832  return (volatile WalIndexHdr*)pWal->apWiData[0];833}834 835/*836** The argument to this macro must be of type u32. On a little-endian837** architecture, it returns the u32 value that results from interpreting838** the 4 bytes as a big-endian value. On a big-endian architecture, it839** returns the value that would be produced by interpreting the 4 bytes840** of the input value as a little-endian integer.841*/842#define BYTESWAP32(x) ( \843    (((x)&0x000000FF)<<24) + (((x)&0x0000FF00)<<8)  \844  + (((x)&0x00FF0000)>>8)  + (((x)&0xFF000000)>>24) \845)846 847/*848** Generate or extend an 8 byte checksum based on the data in849** array aByte[] and the initial values of aIn[0] and aIn[1] (or850** initial values of 0 and 0 if aIn==NULL).851**852** The checksum is written back into aOut[] before returning.853**854** nByte must be a positive multiple of 8.855*/856static void walChecksumBytes(857  int nativeCksum, /* True for native byte-order, false for non-native */858  u8 *a,           /* Content to be checksummed */859  int nByte,       /* Bytes of content in a[].  Must be a multiple of 8. */860  const u32 *aIn,  /* Initial checksum value input */861  u32 *aOut        /* OUT: Final checksum value output */862){863  u32 s1, s2;864  u32 *aData = (u32 *)a;865  u32 *aEnd = (u32 *)&a[nByte];866 867  if( aIn ){868    s1 = aIn[0];869    s2 = aIn[1];870  }else{871    s1 = s2 = 0;872  }873 874  /* nByte is a multiple of 8 between 8 and 65536 */875  assert( nByte>=8 && (nByte&7)==0 && nByte<=65536 );876 877  if( !nativeCksum ){878    do {879      s1 += BYTESWAP32(aData[0]) + s2;880      s2 += BYTESWAP32(aData[1]) + s1;881      aData += 2;882    }while( aData<aEnd );883  }else if( nByte%64==0 ){884    do {885      s1 += *aData++ + s2;886      s2 += *aData++ + s1;887      s1 += *aData++ + s2;888      s2 += *aData++ + s1;889      s1 += *aData++ + s2;890      s2 += *aData++ + s1;891      s1 += *aData++ + s2;892      s2 += *aData++ + s1;893      s1 += *aData++ + s2;894      s2 += *aData++ + s1;895      s1 += *aData++ + s2;896      s2 += *aData++ + s1;897      s1 += *aData++ + s2;898      s2 += *aData++ + s1;899      s1 += *aData++ + s2;900      s2 += *aData++ + s1;901    }while( aData<aEnd );902  }else{903    do {904      s1 += *aData++ + s2;905      s2 += *aData++ + s1;906    }while( aData<aEnd );907  }908  assert( aData==aEnd );909 910  aOut[0] = s1;911  aOut[1] = s2;912}913 914/*915** If there is the possibility of concurrent access to the SHM file916** from multiple threads and/or processes, then do a memory barrier.917*/918static void walShmBarrier(Wal *pWal){919  if( pWal->exclusiveMode!=WAL_HEAPMEMORY_MODE ){920    sqlite3OsShmBarrier(pWal->pDbFd);921  }922}923 924/*925** Add the SQLITE_NO_TSAN as part of the return-type of a function926** definition as a hint that the function contains constructs that927** might give false-positive TSAN warnings.928**929** See tag-20200519-1.930*/931#if defined(__clang__) && !defined(SQLITE_NO_TSAN)932# define SQLITE_NO_TSAN __attribute__((no_sanitize_thread))933#else934# define SQLITE_NO_TSAN935#endif936 937/*938** Write the header information in pWal->hdr into the wal-index.939**940** The checksum on pWal->hdr is updated before it is written.941*/942static SQLITE_NO_TSAN void walIndexWriteHdr(Wal *pWal){943  volatile WalIndexHdr *aHdr = walIndexHdr(pWal);944  const int nCksum = offsetof(WalIndexHdr, aCksum);945 946  assert( pWal->writeLock );947  pWal->hdr.isInit = 1;948  pWal->hdr.iVersion = WALINDEX_MAX_VERSION;949  walChecksumBytes(1, (u8*)&pWal->hdr, nCksum, 0, pWal->hdr.aCksum);950  /* Possible TSAN false-positive.  See tag-20200519-1 */951  memcpy((void*)&aHdr[1], (const void*)&pWal->hdr, sizeof(WalIndexHdr));952  walShmBarrier(pWal);953  memcpy((void*)&aHdr[0], (const void*)&pWal->hdr, sizeof(WalIndexHdr));954}955 956/*957** This function encodes a single frame header and writes it to a buffer958** supplied by the caller. A frame-header is made up of a series of959** 4-byte big-endian integers, as follows:960**961**     0: Page number.962**     4: For commit records, the size of the database image in pages963**        after the commit. For all other records, zero.964**     8: Salt-1 (copied from the wal-header)965**    12: Salt-2 (copied from the wal-header)966**    16: Checksum-1.967**    20: Checksum-2.968*/969static void walEncodeFrame(970  Wal *pWal,                      /* The write-ahead log */971  u32 iPage,                      /* Database page number for frame */972  u32 nTruncate,                  /* New db size (or 0 for non-commit frames) */973  u8 *aData,                      /* Pointer to page data */974  u8 *aFrame                      /* OUT: Write encoded frame here */975){976  int nativeCksum;                /* True for native byte-order checksums */977  u32 *aCksum = pWal->hdr.aFrameCksum;978  assert( WAL_FRAME_HDRSIZE==24 );979  sqlite3Put4byte(&aFrame[0], iPage);980  sqlite3Put4byte(&aFrame[4], nTruncate);981  if( pWal->iReCksum==0 ){982    memcpy(&aFrame[8], pWal->hdr.aSalt, 8);983 984    nativeCksum = (pWal->hdr.bigEndCksum==SQLITE_BIGENDIAN);985    walChecksumBytes(nativeCksum, aFrame, 8, aCksum, aCksum);986    walChecksumBytes(nativeCksum, aData, pWal->szPage, aCksum, aCksum);987 988    sqlite3Put4byte(&aFrame[16], aCksum[0]);989    sqlite3Put4byte(&aFrame[20], aCksum[1]);990  }else{991    memset(&aFrame[8], 0, 16);992  }993}994 995/*996** Check to see if the frame with header in aFrame[] and content997** in aData[] is valid.  If it is a valid frame, fill *piPage and998** *pnTruncate and return true.  Return if the frame is not valid.999*/1000static int walDecodeFrame(1001  Wal *pWal,                      /* The write-ahead log */1002  u32 *piPage,                    /* OUT: Database page number for frame */1003  u32 *pnTruncate,                /* OUT: New db size (or 0 if not commit) */1004  u8 *aData,                      /* Pointer to page data (for checksum) */1005  u8 *aFrame                      /* Frame data */1006){1007  int nativeCksum;                /* True for native byte-order checksums */1008  u32 *aCksum = pWal->hdr.aFrameCksum;1009  u32 pgno;                       /* Page number of the frame */1010  assert( WAL_FRAME_HDRSIZE==24 );1011 1012  /* A frame is only valid if the salt values in the frame-header1013  ** match the salt values in the wal-header.1014  */1015  if( memcmp(&pWal->hdr.aSalt, &aFrame[8], 8)!=0 ){1016    return 0;1017  }1018 1019  /* A frame is only valid if the page number is greater than zero.1020  */1021  pgno = sqlite3Get4byte(&aFrame[0]);1022  if( pgno==0 ){1023    return 0;1024  }1025 1026  /* A frame is only valid if a checksum of the WAL header,1027  ** all prior frames, the first 16 bytes of this frame-header,1028  ** and the frame-data matches the checksum in the last 81029  ** bytes of this frame-header.1030  */1031  nativeCksum = (pWal->hdr.bigEndCksum==SQLITE_BIGENDIAN);1032  walChecksumBytes(nativeCksum, aFrame, 8, aCksum, aCksum);1033  walChecksumBytes(nativeCksum, aData, pWal->szPage, aCksum, aCksum);1034  if( aCksum[0]!=sqlite3Get4byte(&aFrame[16])1035   || aCksum[1]!=sqlite3Get4byte(&aFrame[20])1036  ){1037    /* Checksum failed. */1038    return 0;1039  }1040 1041  /* If we reach this point, the frame is valid.  Return the page number1042  ** and the new database size.1043  */1044  *piPage = pgno;1045  *pnTruncate = sqlite3Get4byte(&aFrame[4]);1046  return 1;1047}1048 1049 1050#if defined(SQLITE_TEST) && defined(SQLITE_DEBUG)1051/*1052** Names of locks.  This routine is used to provide debugging output and is not1053** a part of an ordinary build.1054*/1055static const char *walLockName(int lockIdx){1056  if( lockIdx==WAL_WRITE_LOCK ){1057    return "WRITE-LOCK";1058  }else if( lockIdx==WAL_CKPT_LOCK ){1059    return "CKPT-LOCK";1060  }else if( lockIdx==WAL_RECOVER_LOCK ){1061    return "RECOVER-LOCK";1062  }else{1063    static char zName[15];1064    sqlite3_snprintf(sizeof(zName), zName, "READ-LOCK[%d]",1065                     lockIdx-WAL_READ_LOCK(0));1066    return zName;1067  }1068}1069#endif /*defined(SQLITE_TEST) || defined(SQLITE_DEBUG) */1070 1071 1072/*1073** Set or release locks on the WAL.  Locks are either shared or exclusive.1074** A lock cannot be moved directly between shared and exclusive - it must go1075** through the unlocked state first.1076**1077** In locking_mode=EXCLUSIVE, all of these routines become no-ops.1078*/1079static int walLockShared(Wal *pWal, int lockIdx){1080  int rc;1081  if( pWal->exclusiveMode ) return SQLITE_OK;1082  rc = sqlite3OsShmLock(pWal->pDbFd, lockIdx, 1,1083                        SQLITE_SHM_LOCK | SQLITE_SHM_SHARED);1084  WALTRACE(("WAL%p: acquire SHARED-%s %s\n", pWal,1085            walLockName(lockIdx), rc ? "failed" : "ok"));1086  VVA_ONLY( pWal->lockError = (u8)(rc!=SQLITE_OK && (rc&0xFF)!=SQLITE_BUSY); )1087#ifdef SQLITE_USE_SEH1088  if( rc==SQLITE_OK ) pWal->lockMask |= (1 << lockIdx);1089#endif1090  return rc;1091}1092static void walUnlockShared(Wal *pWal, int lockIdx){1093  if( pWal->exclusiveMode ) return;1094  (void)sqlite3OsShmLock(pWal->pDbFd, lockIdx, 1,1095                         SQLITE_SHM_UNLOCK | SQLITE_SHM_SHARED);1096#ifdef SQLITE_USE_SEH1097  pWal->lockMask &= ~(1 << lockIdx);1098#endif1099  WALTRACE(("WAL%p: release SHARED-%s\n", pWal, walLockName(lockIdx)));1100}1101static int walLockExclusive(Wal *pWal, int lockIdx, int n){1102  int rc;1103  if( pWal->exclusiveMode ) return SQLITE_OK;1104  rc = sqlite3OsShmLock(pWal->pDbFd, lockIdx, n,1105                        SQLITE_SHM_LOCK | SQLITE_SHM_EXCLUSIVE);1106  WALTRACE(("WAL%p: acquire EXCLUSIVE-%s cnt=%d %s\n", pWal,1107            walLockName(lockIdx), n, rc ? "failed" : "ok"));1108  VVA_ONLY( pWal->lockError = (u8)(rc!=SQLITE_OK && (rc&0xFF)!=SQLITE_BUSY); )1109#ifdef SQLITE_USE_SEH1110  if( rc==SQLITE_OK ){1111    pWal->lockMask |= (((1<<n)-1) << (SQLITE_SHM_NLOCK+lockIdx));1112  }1113#endif1114  return rc;1115}1116static void walUnlockExclusive(Wal *pWal, int lockIdx, int n){1117  if( pWal->exclusiveMode ) return;1118  (void)sqlite3OsShmLock(pWal->pDbFd, lockIdx, n,1119                         SQLITE_SHM_UNLOCK | SQLITE_SHM_EXCLUSIVE);1120#ifdef SQLITE_USE_SEH1121  pWal->lockMask &= ~(((1<<n)-1) << (SQLITE_SHM_NLOCK+lockIdx));1122#endif1123  WALTRACE(("WAL%p: release EXCLUSIVE-%s cnt=%d\n", pWal,1124             walLockName(lockIdx), n));1125}1126 1127/*1128** Compute a hash on a page number.  The resulting hash value must land1129** between 0 and (HASHTABLE_NSLOT-1).  The walHashNext() function advances1130** the hash to the next value in the event of a collision.1131*/1132static int walHash(u32 iPage){1133  assert( iPage>0 );1134  assert( (HASHTABLE_NSLOT & (HASHTABLE_NSLOT-1))==0 );1135  return (iPage*HASHTABLE_HASH_1) & (HASHTABLE_NSLOT-1);1136}1137static int walNextHash(int iPriorHash){1138  return (iPriorHash+1)&(HASHTABLE_NSLOT-1);1139}1140 1141/*1142** An instance of the WalHashLoc object is used to describe the location1143** of a page hash table in the wal-index.  This becomes the return value1144** from walHashGet().1145*/1146typedef struct WalHashLoc WalHashLoc;1147struct WalHashLoc {1148  volatile ht_slot *aHash;  /* Start of the wal-index hash table */1149  volatile u32 *aPgno;      /* aPgno[1] is the page of first frame indexed */1150  u32 iZero;                /* One less than the frame number of first indexed*/1151};1152 1153/*1154** Return pointers to the hash table and page number array stored on1155** page iHash of the wal-index. The wal-index is broken into 32KB pages1156** numbered starting from 0.1157**1158** Set output variable pLoc->aHash to point to the start of the hash table1159** in the wal-index file. Set pLoc->iZero to one less than the frame1160** number of the first frame indexed by this hash table. If a1161** slot in the hash table is set to N, it refers to frame number1162** (pLoc->iZero+N) in the log.1163**1164** Finally, set pLoc->aPgno so that pLoc->aPgno[0] is the page number of the1165** first frame indexed by the hash table, frame (pLoc->iZero).1166*/1167static int walHashGet(1168  Wal *pWal,                      /* WAL handle */1169  int iHash,                      /* Find the iHash'th table */1170  WalHashLoc *pLoc                /* OUT: Hash table location */1171){1172  int rc;                         /* Return code */1173 1174  rc = walIndexPage(pWal, iHash, &pLoc->aPgno);1175  assert( rc==SQLITE_OK || iHash>0 );1176 1177  if( pLoc->aPgno ){1178    pLoc->aHash = (volatile ht_slot *)&pLoc->aPgno[HASHTABLE_NPAGE];1179    if( iHash==0 ){1180      pLoc->aPgno = &pLoc->aPgno[WALINDEX_HDR_SIZE/sizeof(u32)];1181      pLoc->iZero = 0;1182    }else{1183      pLoc->iZero = HASHTABLE_NPAGE_ONE + (iHash-1)*HASHTABLE_NPAGE;1184    }1185  }else if( NEVER(rc==SQLITE_OK) ){1186    rc = SQLITE_ERROR;1187  }1188  return rc;1189}1190 1191/*1192** Return the number of the wal-index page that contains the hash-table1193** and page-number array that contain entries corresponding to WAL frame1194** iFrame. The wal-index is broken up into 32KB pages. Wal-index pages1195** are numbered starting from 0.1196*/1197static int walFramePage(u32 iFrame){1198  int iHash = (iFrame+HASHTABLE_NPAGE-HASHTABLE_NPAGE_ONE-1) / HASHTABLE_NPAGE;1199  assert( (iHash==0 || iFrame>HASHTABLE_NPAGE_ONE)1200       && (iHash>=1 || iFrame<=HASHTABLE_NPAGE_ONE)

Showing the first 1,200 of 4622 lines. Download the file for the rest.