AryaWu/sqlite
0
1/*2** 2010 February 13**4** The author disclaims copyright to this source code. In place of5** a legal notice, here is a blessing:6**7** May you do good and not evil.8** May you find forgiveness for yourself and forgive others.9** May you share freely, never taking more than you give.10**11*************************************************************************12**13** This file contains the implementation of a write-ahead log (WAL) used in14** "journal_mode=WAL" mode.15**16** WRITE-AHEAD LOG (WAL) FILE FORMAT17**18** A WAL file consists of a header followed by zero or more "frames".19** Each frame records the revised content of a single page from the20** database file. All changes to the database are recorded by writing21** frames into the WAL. Transactions commit when a frame is written that22** contains a commit marker. A single WAL can and usually does record23** multiple transactions. Periodically, the content of the WAL is24** transferred back into the database file in an operation called a25** "checkpoint".26**27** A single WAL file can be used multiple times. In other words, the28** WAL can fill up with frames and then be checkpointed and then new29** frames can overwrite the old ones. A WAL always grows from beginning30** toward the end. Checksums and counters attached to each frame are31** used to determine which frames within the WAL are valid and which32** are leftovers from prior checkpoints.33**34** The WAL header is 32 bytes in size and consists of the following eight35** big-endian 32-bit unsigned integer values:36**37** 0: Magic number. 0x377f0682 or 0x377f068338** 4: File format version. Currently 300700039** 8: Database page size. Example: 102440** 12: Checkpoint sequence number41** 16: Salt-1, random integer incremented with each checkpoint42** 20: Salt-2, a different random integer changing with each ckpt43** 24: Checksum-1 (first part of checksum for first 24 bytes of header).44** 28: Checksum-2 (second part of checksum for first 24 bytes of header).45**46** Immediately following the wal-header are zero or more frames. Each47** frame consists of a 24-byte frame-header followed by <page-size> bytes48** of page data. The frame-header is six big-endian 32-bit unsigned49** integer values, as follows:50**51** 0: Page number.52** 4: For commit records, the size of the database image in pages53** after the commit. For all other records, zero.54** 8: Salt-1 (copied from the header)55** 12: Salt-2 (copied from the header)56** 16: Checksum-1.57** 20: Checksum-2.58**59** A frame is considered valid if and only if the following conditions are60** true:61**62** (1) The salt-1 and salt-2 values in the frame-header match63** salt values in the wal-header64**65** (2) The checksum values in the final 8 bytes of the frame-header66** exactly match the checksum computed consecutively on the67** WAL header and the first 8 bytes and the content of all frames68** up to and including the current frame.69**70** The checksum is computed using 32-bit big-endian integers if the71** magic number in the first 4 bytes of the WAL is 0x377f0683 and it72** is computed using little-endian if the magic number is 0x377f0682.73** The checksum values are always stored in the frame header in a74** big-endian format regardless of which byte order is used to compute75** the checksum. The checksum is computed by interpreting the input as76** an even number of unsigned 32-bit integers: x[0] through x[N]. The77** algorithm used for the checksum is as follows:78**79** for i from 0 to n-1 step 2:80** s0 += x[i] + s1;81** s1 += x[i+1] + s0;82** endfor83**84** Note that s0 and s1 are both weighted checksums using fibonacci weights85** in reverse order (the largest fibonacci weight occurs on the first element86** of the sequence being summed.) The s1 value spans all 32-bit87** terms of the sequence whereas s0 omits the final term.88**89** On a checkpoint, the WAL is first VFS.xSync-ed, then valid content of the90** WAL is transferred into the database, then the database is VFS.xSync-ed.91** The VFS.xSync operations serve as write barriers - all writes launched92** before the xSync must complete before any write that launches after the93** xSync begins.94**95** After each checkpoint, the salt-1 value is incremented and the salt-296** value is randomized. This prevents old and new frames in the WAL from97** being considered valid at the same time and being checkpointing together98** following a crash.99**100** READER ALGORITHM101**102** To read a page from the database (call it page number P), a reader103** first checks the WAL to see if it contains page P. If so, then the104** last valid instance of page P that is a followed by a commit frame105** or is a commit frame itself becomes the value read. If the WAL106** contains no copies of page P that are valid and which are a commit107** frame or are followed by a commit frame, then page P is read from108** the database file.109**110** To start a read transaction, the reader records the index of the last111** valid frame in the WAL. The reader uses this recorded "mxFrame" value112** for all subsequent read operations. New transactions can be appended113** to the WAL, but as long as the reader uses its original mxFrame value114** and ignores the newly appended content, it will see a consistent snapshot115** of the database from a single point in time. This technique allows116** multiple concurrent readers to view different versions of the database117** content simultaneously.118**119** The reader algorithm in the previous paragraphs works correctly, but120** because frames for page P can appear anywhere within the WAL, the121** reader has to scan the entire WAL looking for page P frames. If the122** WAL is large (multiple megabytes is typical) that scan can be slow,123** and read performance suffers. To overcome this problem, a separate124** data structure called the wal-index is maintained to expedite the125** search for frames of a particular page.126**127** WAL-INDEX FORMAT128**129** Conceptually, the wal-index is shared memory, though VFS implementations130** might choose to implement the wal-index using a mmapped file. Because131** the wal-index is shared memory, SQLite does not support journal_mode=WAL132** on a network filesystem. All users of the database must be able to133** share memory.134**135** In the default unix and windows implementation, the wal-index is a mmapped136** file whose name is the database name with a "-shm" suffix added. For that137** reason, the wal-index is sometimes called the "shm" file.138**139** The wal-index is transient. After a crash, the wal-index can (and should140** be) reconstructed from the original WAL file. In fact, the VFS is required141** to either truncate or zero the header of the wal-index when the last142** connection to it closes. Because the wal-index is transient, it can143** use an architecture-specific format; it does not have to be cross-platform.144** Hence, unlike the database and WAL file formats which store all values145** as big endian, the wal-index can store multi-byte values in the native146** byte order of the host computer.147**148** The purpose of the wal-index is to answer this question quickly: Given149** a page number P and a maximum frame index M, return the index of the150** last frame in the wal before frame M for page P in the WAL, or return151** NULL if there are no frames for page P in the WAL prior to M.152**153** The wal-index consists of a header region, followed by an one or154** more index blocks.155**156** The wal-index header contains the total number of frames within the WAL157** in the mxFrame field.158**159** Each index block except for the first contains information on160** HASHTABLE_NPAGE frames. The first index block contains information on161** HASHTABLE_NPAGE_ONE frames. The values of HASHTABLE_NPAGE_ONE and162** HASHTABLE_NPAGE are selected so that together the wal-index header and163** first index block are the same size as all other index blocks in the164** wal-index. The values are:165**166** HASHTABLE_NPAGE 4096167** HASHTABLE_NPAGE_ONE 4062168**169** Each index block contains two sections, a page-mapping that contains the170** database page number associated with each wal frame, and a hash-table171** that allows readers to query an index block for a specific page number.172** The page-mapping is an array of HASHTABLE_NPAGE (or HASHTABLE_NPAGE_ONE173** for the first index block) 32-bit page numbers. The first entry in the174** first index-block contains the database page number corresponding to the175** first frame in the WAL file. The first entry in the second index block176** in the WAL file corresponds to the (HASHTABLE_NPAGE_ONE+1)th frame in177** the log, and so on.178**179** The last index block in a wal-index usually contains less than the full180** complement of HASHTABLE_NPAGE (or HASHTABLE_NPAGE_ONE) page-numbers,181** depending on the contents of the WAL file. This does not change the182** allocated size of the page-mapping array - the page-mapping array merely183** contains unused entries.184**185** Even without using the hash table, the last frame for page P186** can be found by scanning the page-mapping sections of each index block187** starting with the last index block and moving toward the first, and188** within each index block, starting at the end and moving toward the189** beginning. The first entry that equals P corresponds to the frame190** holding the content for that page.191**192** The hash table consists of HASHTABLE_NSLOT 16-bit unsigned integers.193** HASHTABLE_NSLOT = 2*HASHTABLE_NPAGE, and there is one entry in the194** hash table for each page number in the mapping section, so the hash195** table is never more than half full. The expected number of collisions196** prior to finding a match is 1. Each entry of the hash table is an197** 1-based index of an entry in the mapping section of the same198** index block. Let K be the 1-based index of the largest entry in199** the mapping section. (For index blocks other than the last, K will200** always be exactly HASHTABLE_NPAGE (4096) and for the last index block201** K will be (mxFrame%HASHTABLE_NPAGE).) Unused slots of the hash table202** contain a value of 0.203**204** To look for page P in the hash table, first compute a hash iKey on205** P as follows:206**207** iKey = (P * 383) % HASHTABLE_NSLOT208**209** Then start scanning entries of the hash table, starting with iKey210** (wrapping around to the beginning when the end of the hash table is211** reached) until an unused hash slot is found. Let the first unused slot212** be at index iUnused. (iUnused might be less than iKey if there was213** wrap-around.) Because the hash table is never more than half full,214** the search is guaranteed to eventually hit an unused entry. Let215** iMax be the value between iKey and iUnused, closest to iUnused,216** where aHash[iMax]==P. If there is no iMax entry (if there exists217** no hash slot such that aHash[i]==p) then page P is not in the218** current index block. Otherwise the iMax-th mapping entry of the219** current index block corresponds to the last entry that references220** page P.221**222** A hash search begins with the last index block and moves toward the223** first index block, looking for entries corresponding to page P. On224** average, only two or three slots in each index block need to be225** examined in order to either find the last entry for page P, or to226** establish that no such entry exists in the block. Each index block227** holds over 4000 entries. So two or three index blocks are sufficient228** to cover a typical 10 megabyte WAL file, assuming 1K pages. 8 or 10229** comparisons (on average) suffice to either locate a frame in the230** WAL or to establish that the frame does not exist in the WAL. This231** is much faster than scanning the entire 10MB WAL.232**233** Note that entries are added in order of increasing K. Hence, one234** reader might be using some value K0 and a second reader that started235** at a later time (after additional transactions were added to the WAL236** and to the wal-index) might be using a different value K1, where K1>K0.237** Both readers can use the same hash table and mapping section to get238** the correct result. There may be entries in the hash table with239** K>K0 but to the first reader, those entries will appear to be unused240** slots in the hash table and so the first reader will get an answer as241** if no values greater than K0 had ever been inserted into the hash table242** in the first place - which is what reader one wants. Meanwhile, the243** second reader using K1 will see additional values that were inserted244** later, which is exactly what reader two wants.245**246** When a rollback occurs, the value of K is decreased. Hash table entries247** that correspond to frames greater than the new K value are removed248** from the hash table at this point.249*/250#ifndef SQLITE_OMIT_WAL251 252#include "wal.h"253 254/*255** Trace output macros256*/257#if defined(SQLITE_TEST) && defined(SQLITE_DEBUG)258int sqlite3WalTrace = 0;259# define WALTRACE(X) if(sqlite3WalTrace) sqlite3DebugPrintf X260#else261# define WALTRACE(X)262#endif263 264/*265** The maximum (and only) versions of the wal and wal-index formats266** that may be interpreted by this version of SQLite.267**268** If a client begins recovering a WAL file and finds that (a) the checksum269** values in the wal-header are correct and (b) the version field is not270** WAL_MAX_VERSION, recovery fails and SQLite returns SQLITE_CANTOPEN.271**272** Similarly, if a client successfully reads a wal-index header (i.e. the273** checksum test is successful) and finds that the version field is not274** WALINDEX_MAX_VERSION, then no read-transaction is opened and SQLite275** returns SQLITE_CANTOPEN.276*/277#define WAL_MAX_VERSION 3007000278#define WALINDEX_MAX_VERSION 3007000279 280/*281** Index numbers for various locking bytes. WAL_NREADER is the number282** of available reader locks and should be at least 3. The default283** is SQLITE_SHM_NLOCK==8 and WAL_NREADER==5.284**285** Technically, the various VFSes are free to implement these locks however286** they see fit. However, compatibility is encouraged so that VFSes can287** interoperate. The standard implementation used on both unix and windows288** is for the index number to indicate a byte offset into the289** WalCkptInfo.aLock[] array in the wal-index header. In other words, all290** locks are on the shm file. The WALINDEX_LOCK_OFFSET constant (which291** should be 120) is the location in the shm file for the first locking292** byte.293*/294#define WAL_WRITE_LOCK 0295#define WAL_ALL_BUT_WRITE 1296#define WAL_CKPT_LOCK 1297#define WAL_RECOVER_LOCK 2298#define WAL_READ_LOCK(I) (3+(I))299#define WAL_NREADER (SQLITE_SHM_NLOCK-3)300 301 302/* Object declarations */303typedef struct WalIndexHdr WalIndexHdr;304typedef struct WalIterator WalIterator;305typedef struct WalCkptInfo WalCkptInfo;306 307 308/*309** The following object holds a copy of the wal-index header content.310**311** The actual header in the wal-index consists of two copies of this312** object followed by one instance of the WalCkptInfo object.313** For all versions of SQLite through 3.10.0 and probably beyond,314** the locking bytes (WalCkptInfo.aLock) start at offset 120 and315** the total header size is 136 bytes.316**317** The szPage value can be any power of 2 between 512 and 32768, inclusive.318** Or it can be 1 to represent a 65536-byte page. The latter case was319** added in 3.7.1 when support for 64K pages was added.320*/321struct WalIndexHdr {322 u32 iVersion; /* Wal-index version */323 u32 unused; /* Unused (padding) field */324 u32 iChange; /* Counter incremented each transaction */325 u8 isInit; /* 1 when initialized */326 u8 bigEndCksum; /* True if checksums in WAL are big-endian */327 u16 szPage; /* Database page size in bytes. 1==64K */328 u32 mxFrame; /* Index of last valid frame in the WAL */329 u32 nPage; /* Size of database in pages */330 u32 aFrameCksum[2]; /* Checksum of last frame in log */331 u32 aSalt[2]; /* Two salt values copied from WAL header */332 u32 aCksum[2]; /* Checksum over all prior fields */333};334 335/*336** A copy of the following object occurs in the wal-index immediately337** following the second copy of the WalIndexHdr. This object stores338** information used by checkpoint.339**340** nBackfill is the number of frames in the WAL that have been written341** back into the database. (We call the act of moving content from WAL to342** database "backfilling".) The nBackfill number is never greater than343** WalIndexHdr.mxFrame. nBackfill can only be increased by threads344** holding the WAL_CKPT_LOCK lock (which includes a recovery thread).345** However, a WAL_WRITE_LOCK thread can move the value of nBackfill from346** mxFrame back to zero when the WAL is reset.347**348** nBackfillAttempted is the largest value of nBackfill that a checkpoint349** has attempted to achieve. Normally nBackfill==nBackfillAtempted, however350** the nBackfillAttempted is set before any backfilling is done and the351** nBackfill is only set after all backfilling completes. So if a checkpoint352** crashes, nBackfillAttempted might be larger than nBackfill. The353** WalIndexHdr.mxFrame must never be less than nBackfillAttempted.354**355** The aLock[] field is a set of bytes used for locking. These bytes should356** never be read or written.357**358** There is one entry in aReadMark[] for each reader lock. If a reader359** holds read-lock K, then the value in aReadMark[K] is no greater than360** the mxFrame for that reader. The value READMARK_NOT_USED (0xffffffff)361** for any aReadMark[] means that entry is unused. aReadMark[0] is362** a special case; its value is never used and it exists as a place-holder363** to avoid having to offset aReadMark[] indexes by one. Readers holding364** WAL_READ_LOCK(0) always ignore the entire WAL and read all content365** directly from the database.366**367** The value of aReadMark[K] may only be changed by a thread that368** is holding an exclusive lock on WAL_READ_LOCK(K). Thus, the value of369** aReadMark[K] cannot changed while there is a reader is using that mark370** since the reader will be holding a shared lock on WAL_READ_LOCK(K).371**372** The checkpointer may only transfer frames from WAL to database where373** the frame numbers are less than or equal to every aReadMark[] that is374** in use (that is, every aReadMark[j] for which there is a corresponding375** WAL_READ_LOCK(j)). New readers (usually) pick the aReadMark[] with the376** largest value and will increase an unused aReadMark[] to mxFrame if there377** is not already an aReadMark[] equal to mxFrame. The exception to the378** previous sentence is when nBackfill equals mxFrame (meaning that everything379** in the WAL has been backfilled into the database) then new readers380** will choose aReadMark[0] which has value 0 and hence such reader will381** get all their all content directly from the database file and ignore382** the WAL.383**384** Writers normally append new frames to the end of the WAL. However,385** if nBackfill equals mxFrame (meaning that all WAL content has been386** written back into the database) and if no readers are using the WAL387** (in other words, if there are no WAL_READ_LOCK(i) where i>0) then388** the writer will first "reset" the WAL back to the beginning and start389** writing new content beginning at frame 1.390**391** We assume that 32-bit loads are atomic and so no locks are needed in392** order to read from any aReadMark[] entries.393*/394struct WalCkptInfo {395 u32 nBackfill; /* Number of WAL frames backfilled into DB */396 u32 aReadMark[WAL_NREADER]; /* Reader marks */397 u8 aLock[SQLITE_SHM_NLOCK]; /* Reserved space for locks */398 u32 nBackfillAttempted; /* WAL frames perhaps written, or maybe not */399 u32 notUsed0; /* Available for future enhancements */400};401#define READMARK_NOT_USED 0xffffffff402 403/*404** This is a schematic view of the complete 136-byte header of the405** wal-index file (also known as the -shm file):406**407** +-----------------------------+408** 0: | iVersion | \409** +-----------------------------+ |410** 4: | (unused padding) | |411** +-----------------------------+ |412** 8: | iChange | |413** +-------+-------+-------------+ |414** 12: | bInit | bBig | szPage | |415** +-------+-------+-------------+ |416** 16: | mxFrame | | First copy of the417** +-----------------------------+ | WalIndexHdr object418** 20: | nPage | |419** +-----------------------------+ |420** 24: | aFrameCksum | |421** | | |422** +-----------------------------+ |423** 32: | aSalt | |424** | | |425** +-----------------------------+ |426** 40: | aCksum | |427** | | /428** +-----------------------------+429** 48: | iVersion | \430** +-----------------------------+ |431** 52: | (unused padding) | |432** +-----------------------------+ |433** 56: | iChange | |434** +-------+-------+-------------+ |435** 60: | bInit | bBig | szPage | |436** +-------+-------+-------------+ | Second copy of the437** 64: | mxFrame | | WalIndexHdr438** +-----------------------------+ |439** 68: | nPage | |440** +-----------------------------+ |441** 72: | aFrameCksum | |442** | | |443** +-----------------------------+ |444** 80: | aSalt | |445** | | |446** +-----------------------------+ |447** 88: | aCksum | |448** | | /449** +-----------------------------+450** 96: | nBackfill |451** +-----------------------------+452** 100: | 5 read marks |453** | |454** | |455** | |456** | |457** +-------+-------+------+------+458** 120: | Write | Ckpt | Rcvr | Rd0 | \459** +-------+-------+------+------+ ) 8 lock bytes460** | Read1 | Read2 | Rd3 | Rd4 | /461** +-------+-------+------+------+462** 128: | nBackfillAttempted |463** +-----------------------------+464** 132: | (unused padding) |465** +-----------------------------+466*/467 468/* A block of WALINDEX_LOCK_RESERVED bytes beginning at469** WALINDEX_LOCK_OFFSET is reserved for locks. Since some systems470** only support mandatory file-locks, we do not read or write data471** from the region of the file on which locks are applied.472*/473#define WALINDEX_LOCK_OFFSET (sizeof(WalIndexHdr)*2+offsetof(WalCkptInfo,aLock))474#define WALINDEX_HDR_SIZE (sizeof(WalIndexHdr)*2+sizeof(WalCkptInfo))475 476/* Size of header before each frame in wal */477#define WAL_FRAME_HDRSIZE 24478 479/* Size of write ahead log header, including checksum. */480#define WAL_HDRSIZE 32481 482/* WAL magic value. Either this value, or the same value with the least483** significant bit also set (WAL_MAGIC | 0x00000001) is stored in 32-bit484** big-endian format in the first 4 bytes of a WAL file.485**486** If the LSB is set, then the checksums for each frame within the WAL487** file are calculated by treating all data as an array of 32-bit488** big-endian words. Otherwise, they are calculated by interpreting489** all data as 32-bit little-endian words.490*/491#define WAL_MAGIC 0x377f0682492 493/*494** Return the offset of frame iFrame in the write-ahead log file,495** assuming a database page size of szPage bytes. The offset returned496** is to the start of the write-ahead log frame-header.497*/498#define walFrameOffset(iFrame, szPage) ( \499 WAL_HDRSIZE + ((iFrame)-1)*(i64)((szPage)+WAL_FRAME_HDRSIZE) \500)501 502/*503** An open write-ahead log file is represented by an instance of the504** following object.505**506** writeLock:507** This is usually set to 1 whenever the WRITER lock is held. However,508** if it is set to 2, then the WRITER lock is held but must be released509** by walHandleException() if a SEH exception is thrown.510*/511struct Wal {512 sqlite3_vfs *pVfs; /* The VFS used to create pDbFd */513 sqlite3_file *pDbFd; /* File handle for the database file */514 sqlite3_file *pWalFd; /* File handle for WAL file */515 u32 iCallback; /* Value to pass to log callback (or 0) */516 i64 mxWalSize; /* Truncate WAL to this size upon reset */517 int nWiData; /* Size of array apWiData */518 int szFirstBlock; /* Size of first block written to WAL file */519 volatile u32 **apWiData; /* Pointer to wal-index content in memory */520 u32 szPage; /* Database page size */521 i16 readLock; /* Which read lock is being held. -1 for none */522 u8 syncFlags; /* Flags to use to sync header writes */523 u8 exclusiveMode; /* Non-zero if connection is in exclusive mode */524 u8 writeLock; /* True if in a write transaction */525 u8 ckptLock; /* True if holding a checkpoint lock */526 u8 readOnly; /* WAL_RDWR, WAL_RDONLY, or WAL_SHM_RDONLY */527 u8 truncateOnCommit; /* True to truncate WAL file on commit */528 u8 syncHeader; /* Fsync the WAL header if true */529 u8 padToSectorBoundary; /* Pad transactions out to the next sector */530 u8 bShmUnreliable; /* SHM content is read-only and unreliable */531 WalIndexHdr hdr; /* Wal-index header for current transaction */532 u32 minFrame; /* Ignore wal frames before this one */533 u32 iReCksum; /* On commit, recalculate checksums from here */534 const char *zWalName; /* Name of WAL file */535 u32 nCkpt; /* Checkpoint sequence counter in the wal-header */536#ifdef SQLITE_USE_SEH537 u32 lockMask; /* Mask of locks held */538 void *pFree; /* Pointer to sqlite3_free() if exception thrown */539 u32 *pWiValue; /* Value to write into apWiData[iWiPg] */540 int iWiPg; /* Write pWiValue into apWiData[iWiPg] */541 int iSysErrno; /* System error code following exception */542#endif543#ifdef SQLITE_DEBUG544 int nSehTry; /* Number of nested SEH_TRY{} blocks */545 u8 lockError; /* True if a locking error has occurred */546#endif547#ifdef SQLITE_ENABLE_SNAPSHOT548 WalIndexHdr *pSnapshot; /* Start transaction here if not NULL */549 int bGetSnapshot; /* Transaction opened for sqlite3_get_snapshot() */550#endif551#ifdef SQLITE_ENABLE_SETLK_TIMEOUT552 sqlite3 *db;553#endif554};555 556/*557** Candidate values for Wal.exclusiveMode.558*/559#define WAL_NORMAL_MODE 0560#define WAL_EXCLUSIVE_MODE 1561#define WAL_HEAPMEMORY_MODE 2562 563/*564** Possible values for WAL.readOnly565*/566#define WAL_RDWR 0 /* Normal read/write connection */567#define WAL_RDONLY 1 /* The WAL file is readonly */568#define WAL_SHM_RDONLY 2 /* The SHM file is readonly */569 570/*571** Each page of the wal-index mapping contains a hash-table made up of572** an array of HASHTABLE_NSLOT elements of the following type.573*/574typedef u16 ht_slot;575 576/*577** This structure is used to implement an iterator that loops through578** all frames in the WAL in database page order. Where two or more frames579** correspond to the same database page, the iterator visits only the580** frame most recently written to the WAL (in other words, the frame with581** the largest index).582**583** The internals of this structure are only accessed by:584**585** walIteratorInit() - Create a new iterator,586** walIteratorNext() - Step an iterator,587** walIteratorFree() - Free an iterator.588**589** This functionality is used by the checkpoint code (see walCheckpoint()).590*/591struct WalIterator {592 u32 iPrior; /* Last result returned from the iterator */593 int nSegment; /* Number of entries in aSegment[] */594 struct WalSegment {595 int iNext; /* Next slot in aIndex[] not yet returned */596 ht_slot *aIndex; /* i0, i1, i2... such that aPgno[iN] ascend */597 u32 *aPgno; /* Array of page numbers. */598 int nEntry; /* Nr. of entries in aPgno[] and aIndex[] */599 int iZero; /* Frame number associated with aPgno[0] */600 } aSegment[FLEXARRAY]; /* One for every 32KB page in the wal-index */601};602 603/* Size (in bytes) of a WalIterator object suitable for N or fewer segments */604#define SZ_WALITERATOR(N) \605 (offsetof(WalIterator,aSegment)+(N)*sizeof(struct WalSegment))606 607/*608** Define the parameters of the hash tables in the wal-index file. There609** is a hash-table following every HASHTABLE_NPAGE page numbers in the610** wal-index.611**612** Changing any of these constants will alter the wal-index format and613** create incompatibilities.614*/615#define HASHTABLE_NPAGE 4096 /* Must be power of 2 */616#define HASHTABLE_HASH_1 383 /* Should be prime */617#define HASHTABLE_NSLOT (HASHTABLE_NPAGE*2) /* Must be a power of 2 */618 619/*620** The block of page numbers associated with the first hash-table in a621** wal-index is smaller than usual. This is so that there is a complete622** hash-table on each aligned 32KB page of the wal-index.623*/624#define HASHTABLE_NPAGE_ONE (HASHTABLE_NPAGE - (WALINDEX_HDR_SIZE/sizeof(u32)))625 626/* The wal-index is divided into pages of WALINDEX_PGSZ bytes each. */627#define WALINDEX_PGSZ ( \628 sizeof(ht_slot)*HASHTABLE_NSLOT + HASHTABLE_NPAGE*sizeof(u32) \629)630 631/*632** Structured Exception Handling (SEH) is a Windows-specific technique633** for catching exceptions raised while accessing memory-mapped files.634**635** The -DSQLITE_USE_SEH compile-time option means to use SEH to catch and636** deal with system-level errors that arise during WAL -shm file processing.637** Without this compile-time option, any system-level faults that appear638** while accessing the memory-mapped -shm file will cause a process-wide639** signal to be deliver, which will more than likely cause the entire640** process to exit.641*/642#ifdef SQLITE_USE_SEH643#include <Windows.h>644 645/* Beginning of a block of code in which an exception might occur */646# define SEH_TRY __try { \647 assert( walAssertLockmask(pWal) && pWal->nSehTry==0 ); \648 VVA_ONLY(pWal->nSehTry++);649 650/* The end of a block of code in which an exception might occur */651# define SEH_EXCEPT(X) \652 VVA_ONLY(pWal->nSehTry--); \653 assert( pWal->nSehTry==0 ); \654 } __except( sehExceptionFilter(pWal, GetExceptionCode(), GetExceptionInformation() ) ){ X }655 656/* Simulate a memory-mapping fault in the -shm file for testing purposes */657# define SEH_INJECT_FAULT sehInjectFault(pWal) 658 659/*660** The second argument is the return value of GetExceptionCode() for the 661** current exception. Return EXCEPTION_EXECUTE_HANDLER if the exception code662** indicates that the exception may have been caused by accessing the *-shm 663** file mapping. Or EXCEPTION_CONTINUE_SEARCH otherwise.664*/665static int sehExceptionFilter(Wal *pWal, int eCode, EXCEPTION_POINTERS *p){666 VVA_ONLY(pWal->nSehTry--);667 if( eCode==EXCEPTION_IN_PAGE_ERROR ){668 if( p && p->ExceptionRecord && p->ExceptionRecord->NumberParameters>=3 ){669 /* From MSDN: For this type of exception, the first element of the670 ** ExceptionInformation[] array is a read-write flag - 0 if the exception671 ** was thrown while reading, 1 if while writing. The second element is672 ** the virtual address being accessed. The "third array element specifies673 ** the underlying NTSTATUS code that resulted in the exception". */674 pWal->iSysErrno = (int)p->ExceptionRecord->ExceptionInformation[2];675 }676 return EXCEPTION_EXECUTE_HANDLER;677 }678 return EXCEPTION_CONTINUE_SEARCH;679}680 681/*682** If one is configured, invoke the xTestCallback callback with 650 as683** the argument. If it returns true, throw the same exception that is684** thrown by the system if the *-shm file mapping is accessed after it685** has been invalidated.686*/687static void sehInjectFault(Wal *pWal){688 int res;689 assert( pWal->nSehTry>0 );690 691 res = sqlite3FaultSim(650);692 if( res!=0 ){693 ULONG_PTR aArg[3];694 aArg[0] = 0;695 aArg[1] = 0;696 aArg[2] = (ULONG_PTR)res;697 RaiseException(EXCEPTION_IN_PAGE_ERROR, 0, 3, (const ULONG_PTR*)aArg);698 }699}700 701/*702** There are two ways to use this macro. To set a pointer to be freed703** if an exception is thrown:704**705** SEH_FREE_ON_ERROR(0, pPtr);706**707** and to cancel the same:708**709** SEH_FREE_ON_ERROR(pPtr, 0);710**711** In the first case, there must not already be a pointer registered to712** be freed. In the second case, pPtr must be the registered pointer.713*/714#define SEH_FREE_ON_ERROR(X,Y) \715 assert( (X==0 || Y==0) && pWal->pFree==X ); pWal->pFree = Y716 717/*718** There are two ways to use this macro. To arrange for pWal->apWiData[iPg]719** to be set to pValue if an exception is thrown:720**721** SEH_SET_ON_ERROR(iPg, pValue);722**723** and to cancel the same:724**725** SEH_SET_ON_ERROR(0, 0);726*/727#define SEH_SET_ON_ERROR(X,Y) pWal->iWiPg = X; pWal->pWiValue = Y728 729#else730# define SEH_TRY VVA_ONLY(pWal->nSehTry++);731# define SEH_EXCEPT(X) VVA_ONLY(pWal->nSehTry--); assert( pWal->nSehTry==0 );732# define SEH_INJECT_FAULT assert( pWal->nSehTry>0 );733# define SEH_FREE_ON_ERROR(X,Y)734# define SEH_SET_ON_ERROR(X,Y)735#endif /* ifdef SQLITE_USE_SEH */736 737 738/*739** Obtain a pointer to the iPage'th page of the wal-index. The wal-index740** is broken into pages of WALINDEX_PGSZ bytes. Wal-index pages are741** numbered from zero.742**743** If the wal-index is currently smaller the iPage pages then the size744** of the wal-index might be increased, but only if it is safe to do745** so. It is safe to enlarge the wal-index if pWal->writeLock is true746** or pWal->exclusiveMode==WAL_HEAPMEMORY_MODE.747**748** Three possible result scenarios:749**750** (1) rc==SQLITE_OK and *ppPage==Requested-Wal-Index-Page751** (2) rc>=SQLITE_ERROR and *ppPage==NULL752** (3) rc==SQLITE_OK and *ppPage==NULL // only if iPage==0753**754** Scenario (3) can only occur when pWal->writeLock is false and iPage==0755*/756static SQLITE_NOINLINE int walIndexPageRealloc(757 Wal *pWal, /* The WAL context */758 int iPage, /* The page we seek */759 volatile u32 **ppPage /* Write the page pointer here */760){761 int rc = SQLITE_OK;762 763 /* Enlarge the pWal->apWiData[] array if required */764 if( pWal->nWiData<=iPage ){765 sqlite3_int64 nByte = sizeof(u32*)*(1+(i64)iPage);766 volatile u32 **apNew;767 apNew = (volatile u32 **)sqlite3Realloc((void *)pWal->apWiData, nByte);768 if( !apNew ){769 *ppPage = 0;770 return SQLITE_NOMEM_BKPT;771 }772 memset((void*)&apNew[pWal->nWiData], 0,773 sizeof(u32*)*(iPage+1-pWal->nWiData));774 pWal->apWiData = apNew;775 pWal->nWiData = iPage+1;776 }777 778 /* Request a pointer to the required page from the VFS */779 assert( pWal->apWiData[iPage]==0 );780 if( pWal->exclusiveMode==WAL_HEAPMEMORY_MODE ){781 pWal->apWiData[iPage] = (u32 volatile *)sqlite3MallocZero(WALINDEX_PGSZ);782 if( !pWal->apWiData[iPage] ) rc = SQLITE_NOMEM_BKPT;783 }else{784 rc = sqlite3OsShmMap(pWal->pDbFd, iPage, WALINDEX_PGSZ,785 pWal->writeLock, (void volatile **)&pWal->apWiData[iPage]786 );787 assert( pWal->apWiData[iPage]!=0788 || rc!=SQLITE_OK789 || (pWal->writeLock==0 && iPage==0) );790 testcase( pWal->apWiData[iPage]==0 && rc==SQLITE_OK );791 if( rc==SQLITE_OK ){792 if( iPage>0 && sqlite3FaultSim(600) ) rc = SQLITE_NOMEM;793 }else if( (rc&0xff)==SQLITE_READONLY ){794 pWal->readOnly |= WAL_SHM_RDONLY;795 if( rc==SQLITE_READONLY ){796 rc = SQLITE_OK;797 }798 }799 }800 801 *ppPage = pWal->apWiData[iPage];802 assert( iPage==0 || *ppPage || rc!=SQLITE_OK );803 return rc;804}805static int walIndexPage(806 Wal *pWal, /* The WAL context */807 int iPage, /* The page we seek */808 volatile u32 **ppPage /* Write the page pointer here */809){810 SEH_INJECT_FAULT;811 if( pWal->nWiData<=iPage || (*ppPage = pWal->apWiData[iPage])==0 ){812 return walIndexPageRealloc(pWal, iPage, ppPage);813 }814 return SQLITE_OK;815}816 817/*818** Return a pointer to the WalCkptInfo structure in the wal-index.819*/820static volatile WalCkptInfo *walCkptInfo(Wal *pWal){821 assert( pWal->nWiData>0 && pWal->apWiData[0] );822 SEH_INJECT_FAULT;823 return (volatile WalCkptInfo*)&(pWal->apWiData[0][sizeof(WalIndexHdr)/2]);824}825 826/*827** Return a pointer to the WalIndexHdr structure in the wal-index.828*/829static volatile WalIndexHdr *walIndexHdr(Wal *pWal){830 assert( pWal->nWiData>0 && pWal->apWiData[0] );831 SEH_INJECT_FAULT;832 return (volatile WalIndexHdr*)pWal->apWiData[0];833}834 835/*836** The argument to this macro must be of type u32. On a little-endian837** architecture, it returns the u32 value that results from interpreting838** the 4 bytes as a big-endian value. On a big-endian architecture, it839** returns the value that would be produced by interpreting the 4 bytes840** of the input value as a little-endian integer.841*/842#define BYTESWAP32(x) ( \843 (((x)&0x000000FF)<<24) + (((x)&0x0000FF00)<<8) \844 + (((x)&0x00FF0000)>>8) + (((x)&0xFF000000)>>24) \845)846 847/*848** Generate or extend an 8 byte checksum based on the data in849** array aByte[] and the initial values of aIn[0] and aIn[1] (or850** initial values of 0 and 0 if aIn==NULL).851**852** The checksum is written back into aOut[] before returning.853**854** nByte must be a positive multiple of 8.855*/856static void walChecksumBytes(857 int nativeCksum, /* True for native byte-order, false for non-native */858 u8 *a, /* Content to be checksummed */859 int nByte, /* Bytes of content in a[]. Must be a multiple of 8. */860 const u32 *aIn, /* Initial checksum value input */861 u32 *aOut /* OUT: Final checksum value output */862){863 u32 s1, s2;864 u32 *aData = (u32 *)a;865 u32 *aEnd = (u32 *)&a[nByte];866 867 if( aIn ){868 s1 = aIn[0];869 s2 = aIn[1];870 }else{871 s1 = s2 = 0;872 }873 874 /* nByte is a multiple of 8 between 8 and 65536 */875 assert( nByte>=8 && (nByte&7)==0 && nByte<=65536 );876 877 if( !nativeCksum ){878 do {879 s1 += BYTESWAP32(aData[0]) + s2;880 s2 += BYTESWAP32(aData[1]) + s1;881 aData += 2;882 }while( aData<aEnd );883 }else if( nByte%64==0 ){884 do {885 s1 += *aData++ + s2;886 s2 += *aData++ + s1;887 s1 += *aData++ + s2;888 s2 += *aData++ + s1;889 s1 += *aData++ + s2;890 s2 += *aData++ + s1;891 s1 += *aData++ + s2;892 s2 += *aData++ + s1;893 s1 += *aData++ + s2;894 s2 += *aData++ + s1;895 s1 += *aData++ + s2;896 s2 += *aData++ + s1;897 s1 += *aData++ + s2;898 s2 += *aData++ + s1;899 s1 += *aData++ + s2;900 s2 += *aData++ + s1;901 }while( aData<aEnd );902 }else{903 do {904 s1 += *aData++ + s2;905 s2 += *aData++ + s1;906 }while( aData<aEnd );907 }908 assert( aData==aEnd );909 910 aOut[0] = s1;911 aOut[1] = s2;912}913 914/*915** If there is the possibility of concurrent access to the SHM file916** from multiple threads and/or processes, then do a memory barrier.917*/918static void walShmBarrier(Wal *pWal){919 if( pWal->exclusiveMode!=WAL_HEAPMEMORY_MODE ){920 sqlite3OsShmBarrier(pWal->pDbFd);921 }922}923 924/*925** Add the SQLITE_NO_TSAN as part of the return-type of a function926** definition as a hint that the function contains constructs that927** might give false-positive TSAN warnings.928**929** See tag-20200519-1.930*/931#if defined(__clang__) && !defined(SQLITE_NO_TSAN)932# define SQLITE_NO_TSAN __attribute__((no_sanitize_thread))933#else934# define SQLITE_NO_TSAN935#endif936 937/*938** Write the header information in pWal->hdr into the wal-index.939**940** The checksum on pWal->hdr is updated before it is written.941*/942static SQLITE_NO_TSAN void walIndexWriteHdr(Wal *pWal){943 volatile WalIndexHdr *aHdr = walIndexHdr(pWal);944 const int nCksum = offsetof(WalIndexHdr, aCksum);945 946 assert( pWal->writeLock );947 pWal->hdr.isInit = 1;948 pWal->hdr.iVersion = WALINDEX_MAX_VERSION;949 walChecksumBytes(1, (u8*)&pWal->hdr, nCksum, 0, pWal->hdr.aCksum);950 /* Possible TSAN false-positive. See tag-20200519-1 */951 memcpy((void*)&aHdr[1], (const void*)&pWal->hdr, sizeof(WalIndexHdr));952 walShmBarrier(pWal);953 memcpy((void*)&aHdr[0], (const void*)&pWal->hdr, sizeof(WalIndexHdr));954}955 956/*957** This function encodes a single frame header and writes it to a buffer958** supplied by the caller. A frame-header is made up of a series of959** 4-byte big-endian integers, as follows:960**961** 0: Page number.962** 4: For commit records, the size of the database image in pages963** after the commit. For all other records, zero.964** 8: Salt-1 (copied from the wal-header)965** 12: Salt-2 (copied from the wal-header)966** 16: Checksum-1.967** 20: Checksum-2.968*/969static void walEncodeFrame(970 Wal *pWal, /* The write-ahead log */971 u32 iPage, /* Database page number for frame */972 u32 nTruncate, /* New db size (or 0 for non-commit frames) */973 u8 *aData, /* Pointer to page data */974 u8 *aFrame /* OUT: Write encoded frame here */975){976 int nativeCksum; /* True for native byte-order checksums */977 u32 *aCksum = pWal->hdr.aFrameCksum;978 assert( WAL_FRAME_HDRSIZE==24 );979 sqlite3Put4byte(&aFrame[0], iPage);980 sqlite3Put4byte(&aFrame[4], nTruncate);981 if( pWal->iReCksum==0 ){982 memcpy(&aFrame[8], pWal->hdr.aSalt, 8);983 984 nativeCksum = (pWal->hdr.bigEndCksum==SQLITE_BIGENDIAN);985 walChecksumBytes(nativeCksum, aFrame, 8, aCksum, aCksum);986 walChecksumBytes(nativeCksum, aData, pWal->szPage, aCksum, aCksum);987 988 sqlite3Put4byte(&aFrame[16], aCksum[0]);989 sqlite3Put4byte(&aFrame[20], aCksum[1]);990 }else{991 memset(&aFrame[8], 0, 16);992 }993}994 995/*996** Check to see if the frame with header in aFrame[] and content997** in aData[] is valid. If it is a valid frame, fill *piPage and998** *pnTruncate and return true. Return if the frame is not valid.999*/1000static int walDecodeFrame(1001 Wal *pWal, /* The write-ahead log */1002 u32 *piPage, /* OUT: Database page number for frame */1003 u32 *pnTruncate, /* OUT: New db size (or 0 if not commit) */1004 u8 *aData, /* Pointer to page data (for checksum) */1005 u8 *aFrame /* Frame data */1006){1007 int nativeCksum; /* True for native byte-order checksums */1008 u32 *aCksum = pWal->hdr.aFrameCksum;1009 u32 pgno; /* Page number of the frame */1010 assert( WAL_FRAME_HDRSIZE==24 );1011 1012 /* A frame is only valid if the salt values in the frame-header1013 ** match the salt values in the wal-header.1014 */1015 if( memcmp(&pWal->hdr.aSalt, &aFrame[8], 8)!=0 ){1016 return 0;1017 }1018 1019 /* A frame is only valid if the page number is greater than zero.1020 */1021 pgno = sqlite3Get4byte(&aFrame[0]);1022 if( pgno==0 ){1023 return 0;1024 }1025 1026 /* A frame is only valid if a checksum of the WAL header,1027 ** all prior frames, the first 16 bytes of this frame-header,1028 ** and the frame-data matches the checksum in the last 81029 ** bytes of this frame-header.1030 */1031 nativeCksum = (pWal->hdr.bigEndCksum==SQLITE_BIGENDIAN);1032 walChecksumBytes(nativeCksum, aFrame, 8, aCksum, aCksum);1033 walChecksumBytes(nativeCksum, aData, pWal->szPage, aCksum, aCksum);1034 if( aCksum[0]!=sqlite3Get4byte(&aFrame[16])1035 || aCksum[1]!=sqlite3Get4byte(&aFrame[20])1036 ){1037 /* Checksum failed. */1038 return 0;1039 }1040 1041 /* If we reach this point, the frame is valid. Return the page number1042 ** and the new database size.1043 */1044 *piPage = pgno;1045 *pnTruncate = sqlite3Get4byte(&aFrame[4]);1046 return 1;1047}1048 1049 1050#if defined(SQLITE_TEST) && defined(SQLITE_DEBUG)1051/*1052** Names of locks. This routine is used to provide debugging output and is not1053** a part of an ordinary build.1054*/1055static const char *walLockName(int lockIdx){1056 if( lockIdx==WAL_WRITE_LOCK ){1057 return "WRITE-LOCK";1058 }else if( lockIdx==WAL_CKPT_LOCK ){1059 return "CKPT-LOCK";1060 }else if( lockIdx==WAL_RECOVER_LOCK ){1061 return "RECOVER-LOCK";1062 }else{1063 static char zName[15];1064 sqlite3_snprintf(sizeof(zName), zName, "READ-LOCK[%d]",1065 lockIdx-WAL_READ_LOCK(0));1066 return zName;1067 }1068}1069#endif /*defined(SQLITE_TEST) || defined(SQLITE_DEBUG) */1070 1071 1072/*1073** Set or release locks on the WAL. Locks are either shared or exclusive.1074** A lock cannot be moved directly between shared and exclusive - it must go1075** through the unlocked state first.1076**1077** In locking_mode=EXCLUSIVE, all of these routines become no-ops.1078*/1079static int walLockShared(Wal *pWal, int lockIdx){1080 int rc;1081 if( pWal->exclusiveMode ) return SQLITE_OK;1082 rc = sqlite3OsShmLock(pWal->pDbFd, lockIdx, 1,1083 SQLITE_SHM_LOCK | SQLITE_SHM_SHARED);1084 WALTRACE(("WAL%p: acquire SHARED-%s %s\n", pWal,1085 walLockName(lockIdx), rc ? "failed" : "ok"));1086 VVA_ONLY( pWal->lockError = (u8)(rc!=SQLITE_OK && (rc&0xFF)!=SQLITE_BUSY); )1087#ifdef SQLITE_USE_SEH1088 if( rc==SQLITE_OK ) pWal->lockMask |= (1 << lockIdx);1089#endif1090 return rc;1091}1092static void walUnlockShared(Wal *pWal, int lockIdx){1093 if( pWal->exclusiveMode ) return;1094 (void)sqlite3OsShmLock(pWal->pDbFd, lockIdx, 1,1095 SQLITE_SHM_UNLOCK | SQLITE_SHM_SHARED);1096#ifdef SQLITE_USE_SEH1097 pWal->lockMask &= ~(1 << lockIdx);1098#endif1099 WALTRACE(("WAL%p: release SHARED-%s\n", pWal, walLockName(lockIdx)));1100}1101static int walLockExclusive(Wal *pWal, int lockIdx, int n){1102 int rc;1103 if( pWal->exclusiveMode ) return SQLITE_OK;1104 rc = sqlite3OsShmLock(pWal->pDbFd, lockIdx, n,1105 SQLITE_SHM_LOCK | SQLITE_SHM_EXCLUSIVE);1106 WALTRACE(("WAL%p: acquire EXCLUSIVE-%s cnt=%d %s\n", pWal,1107 walLockName(lockIdx), n, rc ? "failed" : "ok"));1108 VVA_ONLY( pWal->lockError = (u8)(rc!=SQLITE_OK && (rc&0xFF)!=SQLITE_BUSY); )1109#ifdef SQLITE_USE_SEH1110 if( rc==SQLITE_OK ){1111 pWal->lockMask |= (((1<<n)-1) << (SQLITE_SHM_NLOCK+lockIdx));1112 }1113#endif1114 return rc;1115}1116static void walUnlockExclusive(Wal *pWal, int lockIdx, int n){1117 if( pWal->exclusiveMode ) return;1118 (void)sqlite3OsShmLock(pWal->pDbFd, lockIdx, n,1119 SQLITE_SHM_UNLOCK | SQLITE_SHM_EXCLUSIVE);1120#ifdef SQLITE_USE_SEH1121 pWal->lockMask &= ~(((1<<n)-1) << (SQLITE_SHM_NLOCK+lockIdx));1122#endif1123 WALTRACE(("WAL%p: release EXCLUSIVE-%s cnt=%d\n", pWal,1124 walLockName(lockIdx), n));1125}1126 1127/*1128** Compute a hash on a page number. The resulting hash value must land1129** between 0 and (HASHTABLE_NSLOT-1). The walHashNext() function advances1130** the hash to the next value in the event of a collision.1131*/1132static int walHash(u32 iPage){1133 assert( iPage>0 );1134 assert( (HASHTABLE_NSLOT & (HASHTABLE_NSLOT-1))==0 );1135 return (iPage*HASHTABLE_HASH_1) & (HASHTABLE_NSLOT-1);1136}1137static int walNextHash(int iPriorHash){1138 return (iPriorHash+1)&(HASHTABLE_NSLOT-1);1139}1140 1141/*1142** An instance of the WalHashLoc object is used to describe the location1143** of a page hash table in the wal-index. This becomes the return value1144** from walHashGet().1145*/1146typedef struct WalHashLoc WalHashLoc;1147struct WalHashLoc {1148 volatile ht_slot *aHash; /* Start of the wal-index hash table */1149 volatile u32 *aPgno; /* aPgno[1] is the page of first frame indexed */1150 u32 iZero; /* One less than the frame number of first indexed*/1151};1152 1153/*1154** Return pointers to the hash table and page number array stored on1155** page iHash of the wal-index. The wal-index is broken into 32KB pages1156** numbered starting from 0.1157**1158** Set output variable pLoc->aHash to point to the start of the hash table1159** in the wal-index file. Set pLoc->iZero to one less than the frame1160** number of the first frame indexed by this hash table. If a1161** slot in the hash table is set to N, it refers to frame number1162** (pLoc->iZero+N) in the log.1163**1164** Finally, set pLoc->aPgno so that pLoc->aPgno[0] is the page number of the1165** first frame indexed by the hash table, frame (pLoc->iZero).1166*/1167static int walHashGet(1168 Wal *pWal, /* WAL handle */1169 int iHash, /* Find the iHash'th table */1170 WalHashLoc *pLoc /* OUT: Hash table location */1171){1172 int rc; /* Return code */1173 1174 rc = walIndexPage(pWal, iHash, &pLoc->aPgno);1175 assert( rc==SQLITE_OK || iHash>0 );1176 1177 if( pLoc->aPgno ){1178 pLoc->aHash = (volatile ht_slot *)&pLoc->aPgno[HASHTABLE_NPAGE];1179 if( iHash==0 ){1180 pLoc->aPgno = &pLoc->aPgno[WALINDEX_HDR_SIZE/sizeof(u32)];1181 pLoc->iZero = 0;1182 }else{1183 pLoc->iZero = HASHTABLE_NPAGE_ONE + (iHash-1)*HASHTABLE_NPAGE;1184 }1185 }else if( NEVER(rc==SQLITE_OK) ){1186 rc = SQLITE_ERROR;1187 }1188 return rc;1189}1190 1191/*1192** Return the number of the wal-index page that contains the hash-table1193** and page-number array that contain entries corresponding to WAL frame1194** iFrame. The wal-index is broken up into 32KB pages. Wal-index pages1195** are numbered starting from 0.1196*/1197static int walFramePage(u32 iFrame){1198 int iHash = (iFrame+HASHTABLE_NPAGE-HASHTABLE_NPAGE_ONE-1) / HASHTABLE_NPAGE;1199 assert( (iHash==0 || iFrame>HASHTABLE_NPAGE_ONE)1200 && (iHash>=1 || iFrame<=HASHTABLE_NPAGE_ONE)