AryaWu/sqlite
0
1/*2** 2001 September 153**4** The author disclaims copyright to this source code. In place of5** a legal notice, here is a blessing:6**7** May you do good and not evil.8** May you find forgiveness for yourself and forgive others.9** May you share freely, never taking more than you give.10**11*************************************************************************12** This is the implementation of the page cache subsystem or "pager".13**14** The pager is used to access a database disk file. It implements15** atomic commit and rollback through the use of a journal file that16** is separate from the database file. The pager also implements file17** locking to prevent two processes from writing the same database18** file simultaneously, or one process from reading the database while19** another is writing.20*/21#ifndef SQLITE_OMIT_DISKIO22#include "sqliteInt.h"23#include "wal.h"24 25 26/******************* NOTES ON THE DESIGN OF THE PAGER ************************27**28** This comment block describes invariants that hold when using a rollback29** journal. These invariants do not apply for journal_mode=WAL,30** journal_mode=MEMORY, or journal_mode=OFF.31**32** Within this comment block, a page is deemed to have been synced33** automatically as soon as it is written when PRAGMA synchronous=OFF.34** Otherwise, the page is not synced until the xSync method of the VFS35** is called successfully on the file containing the page.36**37** Definition: A page of the database file is said to be "overwriteable" if38** one or more of the following are true about the page:39**40** (a) The original content of the page as it was at the beginning of41** the transaction has been written into the rollback journal and42** synced.43**44** (b) The page was a freelist leaf page at the start of the transaction.45**46** (c) The page number is greater than the largest page that existed in47** the database file at the start of the transaction.48**49** (1) A page of the database file is never overwritten unless one of the50** following are true:51**52** (a) The page and all other pages on the same sector are overwriteable.53**54** (b) The atomic page write optimization is enabled, and the entire55** transaction other than the update of the transaction sequence56** number consists of a single page change.57**58** (2) The content of a page written into the rollback journal exactly matches59** both the content in the database when the rollback journal was written60** and the content in the database at the beginning of the current61** transaction.62**63** (3) Writes to the database file are an integer multiple of the page size64** in length and are aligned on a page boundary.65**66** (4) Reads from the database file are either aligned on a page boundary and67** an integer multiple of the page size in length or are taken from the68** first 100 bytes of the database file.69**70** (5) All writes to the database file are synced prior to the rollback journal71** being deleted, truncated, or zeroed.72**73** (6) If a super-journal file is used, then all writes to the database file74** are synced prior to the super-journal being deleted.75**76** Definition: Two databases (or the same database at two points it time)77** are said to be "logically equivalent" if they give the same answer to78** all queries. Note in particular the content of freelist leaf79** pages can be changed arbitrarily without affecting the logical equivalence80** of the database.81**82** (7) At any time, if any subset, including the empty set and the total set,83** of the unsynced changes to a rollback journal are removed and the84** journal is rolled back, the resulting database file will be logically85** equivalent to the database file at the beginning of the transaction.86**87** (8) When a transaction is rolled back, the xTruncate method of the VFS88** is called to restore the database file to the same size it was at89** the beginning of the transaction. (In some VFSes, the xTruncate90** method is a no-op, but that does not change the fact the SQLite will91** invoke it.)92**93** (9) Whenever the database file is modified, at least one bit in the range94** of bytes from 24 through 39 inclusive will be changed prior to releasing95** the EXCLUSIVE lock, thus signaling other connections on the same96** database to flush their caches.97**98** (10) The pattern of bits in bytes 24 through 39 shall not repeat in less99** than one billion transactions.100**101** (11) A database file is well-formed at the beginning and at the conclusion102** of every transaction.103**104** (12) An EXCLUSIVE lock is held on the database file when writing to105** the database file.106**107** (13) A SHARED lock is held on the database file while reading any108** content out of the database file.109**110******************************************************************************/111 112/*113** Macros for troubleshooting. Normally turned off114*/115#if 0116int sqlite3PagerTrace=1; /* True to enable tracing */117#define sqlite3DebugPrintf printf118#define PAGERTRACE(X) if( sqlite3PagerTrace ){ sqlite3DebugPrintf X; }119#else120#define PAGERTRACE(X)121#endif122 123/*124** The following two macros are used within the PAGERTRACE() macros above125** to print out file-descriptors.126**127** PAGERID() takes a pointer to a Pager struct as its argument. The128** associated file-descriptor is returned. FILEHANDLEID() takes an sqlite3_file129** struct as its argument.130*/131#define PAGERID(p) (SQLITE_PTR_TO_INT(p->fd))132#define FILEHANDLEID(fd) (SQLITE_PTR_TO_INT(fd))133 134/*135** The Pager.eState variable stores the current 'state' of a pager. A136** pager may be in any one of the seven states shown in the following137** state diagram.138**139** OPEN <------+------+140** | | |141** V | |142** +---------> READER-------+ |143** | | |144** | V |145** |<-------WRITER_LOCKED------> ERROR146** | | ^ 147** | V |148** |<------WRITER_CACHEMOD-------->|149** | | |150** | V |151** |<-------WRITER_DBMOD---------->|152** | | |153** | V |154** +<------WRITER_FINISHED-------->+155**156**157** List of state transitions and the C [function] that performs each:158**159** OPEN -> READER [sqlite3PagerSharedLock]160** READER -> OPEN [pager_unlock]161**162** READER -> WRITER_LOCKED [sqlite3PagerBegin]163** WRITER_LOCKED -> WRITER_CACHEMOD [pager_open_journal]164** WRITER_CACHEMOD -> WRITER_DBMOD [syncJournal]165** WRITER_DBMOD -> WRITER_FINISHED [sqlite3PagerCommitPhaseOne]166** WRITER_*** -> READER [pager_end_transaction]167**168** WRITER_*** -> ERROR [pager_error]169** ERROR -> OPEN [pager_unlock]170**171**172** OPEN:173**174** The pager starts up in this state. Nothing is guaranteed in this175** state - the file may or may not be locked and the database size is176** unknown. The database may not be read or written.177**178** * No read or write transaction is active.179** * Any lock, or no lock at all, may be held on the database file.180** * The dbSize, dbOrigSize and dbFileSize variables may not be trusted.181**182** READER:183**184** In this state all the requirements for reading the database in185** rollback (non-WAL) mode are met. Unless the pager is (or recently186** was) in exclusive-locking mode, a user-level read transaction is187** open. The database size is known in this state.188**189** A connection running with locking_mode=normal enters this state when190** it opens a read-transaction on the database and returns to state191** OPEN after the read-transaction is completed. However a connection192** running in locking_mode=exclusive (including temp databases) remains in193** this state even after the read-transaction is closed. The only way194** a locking_mode=exclusive connection can transition from READER to OPEN195** is via the ERROR state (see below).196**197** * A read transaction may be active (but a write-transaction cannot).198** * A SHARED or greater lock is held on the database file.199** * The dbSize variable may be trusted (even if a user-level read200** transaction is not active). The dbOrigSize and dbFileSize variables201** may not be trusted at this point.202** * If the database is a WAL database, then the WAL connection is open.203** * Even if a read-transaction is not open, it is guaranteed that204** there is no hot-journal in the file-system.205**206** WRITER_LOCKED:207**208** The pager moves to this state from READER when a write-transaction209** is first opened on the database. In WRITER_LOCKED state, all locks210** required to start a write-transaction are held, but no actual211** modifications to the cache or database have taken place.212**213** In rollback mode, a RESERVED or (if the transaction was opened with214** BEGIN EXCLUSIVE) EXCLUSIVE lock is obtained on the database file when215** moving to this state, but the journal file is not written to or opened216** to in this state. If the transaction is committed or rolled back while217** in WRITER_LOCKED state, all that is required is to unlock the database218** file.219**220** IN WAL mode, WalBeginWriteTransaction() is called to lock the log file.221** If the connection is running with locking_mode=exclusive, an attempt222** is made to obtain an EXCLUSIVE lock on the database file.223**224** * A write transaction is active.225** * If the connection is open in rollback-mode, a RESERVED or greater226** lock is held on the database file.227** * If the connection is open in WAL-mode, a WAL write transaction228** is open (i.e. sqlite3WalBeginWriteTransaction() has been successfully229** called).230** * The dbSize, dbOrigSize and dbFileSize variables are all valid.231** * The contents of the pager cache have not been modified.232** * The journal file may or may not be open.233** * Nothing (not even the first header) has been written to the journal.234**235** WRITER_CACHEMOD:236**237** A pager moves from WRITER_LOCKED state to this state when a page is238** first modified by the upper layer. In rollback mode the journal file239** is opened (if it is not already open) and a header written to the240** start of it. The database file on disk has not been modified.241**242** * A write transaction is active.243** * A RESERVED or greater lock is held on the database file.244** * The journal file is open and the first header has been written245** to it, but the header has not been synced to disk.246** * The contents of the page cache have been modified.247**248** WRITER_DBMOD:249**250** The pager transitions from WRITER_CACHEMOD into WRITER_DBMOD state251** when it modifies the contents of the database file. WAL connections252** never enter this state (since they do not modify the database file,253** just the log file).254**255** * A write transaction is active.256** * An EXCLUSIVE or greater lock is held on the database file.257** * The journal file is open and the first header has been written258** and synced to disk.259** * The contents of the page cache have been modified (and possibly260** written to disk).261**262** WRITER_FINISHED:263**264** It is not possible for a WAL connection to enter this state.265**266** A rollback-mode pager changes to WRITER_FINISHED state from WRITER_DBMOD267** state after the entire transaction has been successfully written into the268** database file. In this state the transaction may be committed simply269** by finalizing the journal file. Once in WRITER_FINISHED state, it is270** not possible to modify the database further. At this point, the upper271** layer must either commit or rollback the transaction.272**273** * A write transaction is active.274** * An EXCLUSIVE or greater lock is held on the database file.275** * All writing and syncing of journal and database data has finished.276** If no error occurred, all that remains is to finalize the journal to277** commit the transaction. If an error did occur, the caller will need278** to rollback the transaction.279**280** ERROR:281**282** The ERROR state is entered when an IO or disk-full error (including283** SQLITE_IOERR_NOMEM) occurs at a point in the code that makes it284** difficult to be sure that the in-memory pager state (cache contents,285** db size etc.) are consistent with the contents of the file-system.286**287** Temporary pager files may enter the ERROR state, but in-memory pagers288** cannot.289**290** For example, if an IO error occurs while performing a rollback,291** the contents of the page-cache may be left in an inconsistent state.292** At this point it would be dangerous to change back to READER state293** (as usually happens after a rollback). Any subsequent readers might294** report database corruption (due to the inconsistent cache), and if295** they upgrade to writers, they may inadvertently corrupt the database296** file. To avoid this hazard, the pager switches into the ERROR state297** instead of READER following such an error.298**299** Once it has entered the ERROR state, any attempt to use the pager300** to read or write data returns an error. Eventually, once all301** outstanding transactions have been abandoned, the pager is able to302** transition back to OPEN state, discarding the contents of the303** page-cache and any other in-memory state at the same time. Everything304** is reloaded from disk (and, if necessary, hot-journal rollback performed)305** when a read-transaction is next opened on the pager (transitioning306** the pager into READER state). At that point the system has recovered307** from the error.308**309** Specifically, the pager jumps into the ERROR state if:310**311** 1. An error occurs while attempting a rollback. This happens in312** function sqlite3PagerRollback().313**314** 2. An error occurs while attempting to finalize a journal file315** following a commit in function sqlite3PagerCommitPhaseTwo().316**317** 3. An error occurs while attempting to write to the journal or318** database file in function pagerStress() in order to free up319** memory.320**321** In other cases, the error is returned to the b-tree layer. The b-tree322** layer then attempts a rollback operation. If the error condition323** persists, the pager enters the ERROR state via condition (1) above.324**325** Condition (3) is necessary because it can be triggered by a read-only326** statement executed within a transaction. In this case, if the error327** code were simply returned to the user, the b-tree layer would not328** automatically attempt a rollback, as it assumes that an error in a329** read-only statement cannot leave the pager in an internally inconsistent330** state.331**332** * The Pager.errCode variable is set to something other than SQLITE_OK.333** * There are one or more outstanding references to pages (after the334** last reference is dropped the pager should move back to OPEN state).335** * The pager is not an in-memory pager.336** 337**338** Notes:339**340** * A pager is never in WRITER_DBMOD or WRITER_FINISHED state if the341** connection is open in WAL mode. A WAL connection is always in one342** of the first four states.343**344** * Normally, a connection open in exclusive mode is never in PAGER_OPEN345** state. There are two exceptions: immediately after exclusive-mode has346** been turned on (and before any read or write transactions are347** executed), and when the pager is leaving the "error state".348**349** * See also: assert_pager_state().350*/351#define PAGER_OPEN 0352#define PAGER_READER 1353#define PAGER_WRITER_LOCKED 2354#define PAGER_WRITER_CACHEMOD 3355#define PAGER_WRITER_DBMOD 4356#define PAGER_WRITER_FINISHED 5357#define PAGER_ERROR 6358 359/*360** The Pager.eLock variable is almost always set to one of the361** following locking-states, according to the lock currently held on362** the database file: NO_LOCK, SHARED_LOCK, RESERVED_LOCK or EXCLUSIVE_LOCK.363** This variable is kept up to date as locks are taken and released by364** the pagerLockDb() and pagerUnlockDb() wrappers.365**366** If the VFS xLock() or xUnlock() returns an error other than SQLITE_BUSY367** (i.e. one of the SQLITE_IOERR subtypes), it is not clear whether or not368** the operation was successful. In these circumstances pagerLockDb() and369** pagerUnlockDb() take a conservative approach - eLock is always updated370** when unlocking the file, and only updated when locking the file if the371** VFS call is successful. This way, the Pager.eLock variable may be set372** to a less exclusive (lower) value than the lock that is actually held373** at the system level, but it is never set to a more exclusive value.374**375** This is usually safe. If an xUnlock fails or appears to fail, there may376** be a few redundant xLock() calls or a lock may be held for longer than377** required, but nothing really goes wrong.378**379** The exception is when the database file is unlocked as the pager moves380** from ERROR to OPEN state. At this point there may be a hot-journal file381** in the file-system that needs to be rolled back (as part of an OPEN->SHARED382** transition, by the same pager or any other). If the call to xUnlock()383** fails at this point and the pager is left holding an EXCLUSIVE lock, this384** can confuse the call to xCheckReservedLock() call made later as part385** of hot-journal detection.386**387** xCheckReservedLock() is defined as returning true "if there is a RESERVED388** lock held by this process or any others". So xCheckReservedLock may389** return true because the caller itself is holding an EXCLUSIVE lock (but390** doesn't know it because of a previous error in xUnlock). If this happens391** a hot-journal may be mistaken for a journal being created by an active392** transaction in another process, causing SQLite to read from the database393** without rolling it back.394**395** To work around this, if a call to xUnlock() fails when unlocking the396** database in the ERROR state, Pager.eLock is set to UNKNOWN_LOCK. It397** is only changed back to a real locking state after a successful call398** to xLock(EXCLUSIVE). Also, the code to do the OPEN->SHARED state transition399** omits the check for a hot-journal if Pager.eLock is set to UNKNOWN_LOCK400** lock. Instead, it assumes a hot-journal exists and obtains an EXCLUSIVE401** lock on the database file before attempting to roll it back. See function402** PagerSharedLock() for more detail.403**404** Pager.eLock may only be set to UNKNOWN_LOCK when the pager is in405** PAGER_OPEN state.406*/407#define UNKNOWN_LOCK (EXCLUSIVE_LOCK+1)408 409/*410** The maximum allowed sector size. 64KiB. If the xSectorsize() method411** returns a value larger than this, then MAX_SECTOR_SIZE is used instead.412** This could conceivably cause corruption following a power failure on413** such a system. This is currently an undocumented limit.414*/415#define MAX_SECTOR_SIZE 0x10000416 417 418/*419** An instance of the following structure is allocated for each active420** savepoint and statement transaction in the system. All such structures421** are stored in the Pager.aSavepoint[] array, which is allocated and422** resized using sqlite3Realloc().423**424** When a savepoint is created, the PagerSavepoint.iHdrOffset field is425** set to 0. If a journal-header is written into the main journal while426** the savepoint is active, then iHdrOffset is set to the byte offset427** immediately following the last journal record written into the main428** journal before the journal-header. This is required during savepoint429** rollback (see pagerPlaybackSavepoint()).430*/431typedef struct PagerSavepoint PagerSavepoint;432struct PagerSavepoint {433 i64 iOffset; /* Starting offset in main journal */434 i64 iHdrOffset; /* See above */435 Bitvec *pInSavepoint; /* Set of pages in this savepoint */436 Pgno nOrig; /* Original number of pages in file */437 Pgno iSubRec; /* Index of first record in sub-journal */438 int bTruncateOnRelease; /* If stmt journal may be truncated on RELEASE */439#ifndef SQLITE_OMIT_WAL440 u32 aWalData[WAL_SAVEPOINT_NDATA]; /* WAL savepoint context */441#endif442};443 444/*445** Bits of the Pager.doNotSpill flag. See further description below.446*/447#define SPILLFLAG_OFF 0x01 /* Never spill cache. Set via pragma */448#define SPILLFLAG_ROLLBACK 0x02 /* Current rolling back, so do not spill */449#define SPILLFLAG_NOSYNC 0x04 /* Spill is ok, but do not sync */450 451/*452** An open page cache is an instance of struct Pager. A description of453** some of the more important member variables follows:454**455** eState456**457** The current 'state' of the pager object. See the comment and state458** diagram above for a description of the pager state.459**460** eLock461**462** For a real on-disk database, the current lock held on the database file -463** NO_LOCK, SHARED_LOCK, RESERVED_LOCK or EXCLUSIVE_LOCK.464**465** For a temporary or in-memory database (neither of which require any466** locks), this variable is always set to EXCLUSIVE_LOCK. Since such467** databases always have Pager.exclusiveMode==1, this tricks the pager468** logic into thinking that it already has all the locks it will ever469** need (and no reason to release them).470**471** In some (obscure) circumstances, this variable may also be set to472** UNKNOWN_LOCK. See the comment above the #define of UNKNOWN_LOCK for473** details.474**475** changeCountDone476**477** This boolean variable is used to make sure that the change-counter478** (the 4-byte header field at byte offset 24 of the database file) is479** not updated more often than necessary.480**481** It is set to true when the change-counter field is updated, which482** can only happen if an exclusive lock is held on the database file.483** It is cleared (set to false) whenever an exclusive lock is484** relinquished on the database file. Each time a transaction is committed,485** The changeCountDone flag is inspected. If it is true, the work of486** updating the change-counter is omitted for the current transaction.487**488** This mechanism means that when running in exclusive mode, a connection489** need only update the change-counter once, for the first transaction490** committed.491**492** setSuper493**494** When PagerCommitPhaseOne() is called to commit a transaction, it may495** (or may not) specify a super-journal name to be written into the496** journal file before it is synced to disk.497**498** Whether or not a journal file contains a super-journal pointer affects499** the way in which the journal file is finalized after the transaction is500** committed or rolled back when running in "journal_mode=PERSIST" mode.501** If a journal file does not contain a super-journal pointer, it is502** finalized by overwriting the first journal header with zeroes. If503** it does contain a super-journal pointer the journal file is finalized504** by truncating it to zero bytes, just as if the connection were505** running in "journal_mode=truncate" mode.506**507** Journal files that contain super-journal pointers cannot be finalized508** simply by overwriting the first journal-header with zeroes, as the509** super-journal pointer could interfere with hot-journal rollback of any510** subsequently interrupted transaction that reuses the journal file.511**512** The flag is cleared as soon as the journal file is finalized (either513** by PagerCommitPhaseTwo or PagerRollback). If an IO error prevents the514** journal file from being successfully finalized, the setSuper flag515** is cleared anyway (and the pager will move to ERROR state).516**517** doNotSpill518**519** This variables control the behavior of cache-spills (calls made by520** the pcache module to the pagerStress() routine to write cached data521** to the file-system in order to free up memory).522**523** When bits SPILLFLAG_OFF or SPILLFLAG_ROLLBACK of doNotSpill are set,524** writing to the database from pagerStress() is disabled altogether.525** The SPILLFLAG_ROLLBACK case is done in a very obscure case that526** comes up during savepoint rollback that requires the pcache module527** to allocate a new page to prevent the journal file from being written528** while it is being traversed by code in pager_playback(). The SPILLFLAG_OFF529** case is a user preference.530**531** If the SPILLFLAG_NOSYNC bit is set, writing to the database from532** pagerStress() is permitted, but syncing the journal file is not.533** This flag is set by sqlite3PagerWrite() when the file-system sector-size534** is larger than the database page-size in order to prevent a journal sync535** from happening in between the journalling of two pages on the same sector.536**537** subjInMemory538**539** This is a boolean variable. If true, then any required sub-journal540** is opened as an in-memory journal file. If false, then in-memory541** sub-journals are only used for in-memory pager files.542**543** This variable is updated by the upper layer each time a new544** write-transaction is opened.545**546** dbSize, dbOrigSize, dbFileSize547**548** Variable dbSize is set to the number of pages in the database file.549** It is valid in PAGER_READER and higher states (all states except for550** OPEN and ERROR).551**552** dbSize is set based on the size of the database file, which may be553** larger than the size of the database (the value stored at offset554** 28 of the database header by the btree). If the size of the file555** is not an integer multiple of the page-size, the value stored in556** dbSize is rounded down (i.e. a 5KB file with 2K page-size has dbSize==2).557** Except, any file that is greater than 0 bytes in size is considered558** to have at least one page. (i.e. a 1KB file with 2K page-size leads559** to dbSize==1).560**561** During a write-transaction, if pages with page-numbers greater than562** dbSize are modified in the cache, dbSize is updated accordingly.563** Similarly, if the database is truncated using PagerTruncateImage(),564** dbSize is updated.565**566** Variables dbOrigSize and dbFileSize are valid in states567** PAGER_WRITER_LOCKED and higher. dbOrigSize is a copy of the dbSize568** variable at the start of the transaction. It is used during rollback,569** and to determine whether or not pages need to be journalled before570** being modified.571**572** Throughout a write-transaction, dbFileSize contains the size of573** the file on disk in pages. It is set to a copy of dbSize when the574** write-transaction is first opened, and updated when VFS calls are made575** to write or truncate the database file on disk.576**577** The only reason the dbFileSize variable is required is to suppress578** unnecessary calls to xTruncate() after committing a transaction. If,579** when a transaction is committed, the dbFileSize variable indicates580** that the database file is larger than the database image (Pager.dbSize),581** pager_truncate() is called. The pager_truncate() call uses xFilesize()582** to measure the database file on disk, and then truncates it if required.583** dbFileSize is not used when rolling back a transaction. In this case584** pager_truncate() is called unconditionally (which means there may be585** a call to xFilesize() that is not strictly required). In either case,586** pager_truncate() may cause the file to become smaller or larger.587**588** dbHintSize589**590** The dbHintSize variable is used to limit the number of calls made to591** the VFS xFileControl(FCNTL_SIZE_HINT) method.592**593** dbHintSize is set to a copy of the dbSize variable when a594** write-transaction is opened (at the same time as dbFileSize and595** dbOrigSize). If the xFileControl(FCNTL_SIZE_HINT) method is called,596** dbHintSize is increased to the number of pages that correspond to the597** size-hint passed to the method call. See pager_write_pagelist() for598** details.599**600** errCode601**602** The Pager.errCode variable is only ever used in PAGER_ERROR state. It603** is set to zero in all other states. In PAGER_ERROR state, Pager.errCode604** is always set to SQLITE_FULL, SQLITE_IOERR or one of the SQLITE_IOERR_XXX605** sub-codes.606**607** syncFlags, walSyncFlags608**609** syncFlags is either SQLITE_SYNC_NORMAL (0x02) or SQLITE_SYNC_FULL (0x03).610** syncFlags is used for rollback mode. walSyncFlags is used for WAL mode611** and contains the flags used to sync the checkpoint operations in the612** lower two bits, and sync flags used for transaction commits in the WAL613** file in bits 0x04 and 0x08. In other words, to get the correct sync flags614** for checkpoint operations, use (walSyncFlags&0x03) and to get the correct615** sync flags for transaction commit, use ((walSyncFlags>>2)&0x03). Note616** that with synchronous=NORMAL in WAL mode, transaction commit is not synced617** meaning that the 0x04 and 0x08 bits are both zero.618*/619struct Pager {620 sqlite3_vfs *pVfs; /* OS functions to use for IO */621 u8 exclusiveMode; /* Boolean. True if locking_mode==EXCLUSIVE */622 u8 journalMode; /* One of the PAGER_JOURNALMODE_* values */623 u8 useJournal; /* Use a rollback journal on this file */624 u8 noSync; /* Do not sync the journal if true */625 u8 fullSync; /* Do extra syncs of the journal for robustness */626 u8 extraSync; /* sync directory after journal delete */627 u8 syncFlags; /* SYNC_NORMAL or SYNC_FULL otherwise */628 u8 walSyncFlags; /* See description above */629 u8 tempFile; /* zFilename is a temporary or immutable file */630 u8 noLock; /* Do not lock (except in WAL mode) */631 u8 readOnly; /* True for a read-only database */632 u8 memDb; /* True to inhibit all file I/O */633 u8 memVfs; /* VFS-implemented memory database */634 635 /**************************************************************************636 ** The following block contains those class members that change during637 ** routine operation. Class members not in this block are either fixed638 ** when the pager is first created or else only change when there is a639 ** significant mode change (such as changing the page_size, locking_mode,640 ** or the journal_mode). From another view, these class members describe641 ** the "state" of the pager, while other class members describe the642 ** "configuration" of the pager.643 */644 u8 eState; /* Pager state (OPEN, READER, WRITER_LOCKED..) */645 u8 eLock; /* Current lock held on database file */646 u8 changeCountDone; /* Set after incrementing the change-counter */647 u8 setSuper; /* Super-jrnl name is written into jrnl */648 u8 doNotSpill; /* Do not spill the cache when non-zero */649 u8 subjInMemory; /* True to use in-memory sub-journals */650 u8 bUseFetch; /* True to use xFetch() */651 u8 hasHeldSharedLock; /* True if a shared lock has ever been held */652 Pgno dbSize; /* Number of pages in the database */653 Pgno dbOrigSize; /* dbSize before the current transaction */654 Pgno dbFileSize; /* Number of pages in the database file */655 Pgno dbHintSize; /* Value passed to FCNTL_SIZE_HINT call */656 int errCode; /* One of several kinds of errors */657 int nRec; /* Pages journalled since last j-header written */658 u32 cksumInit; /* Quasi-random value added to every checksum */659 u32 nSubRec; /* Number of records written to sub-journal */660 Bitvec *pInJournal; /* One bit for each page in the database file */661 sqlite3_file *fd; /* File descriptor for database */662 sqlite3_file *jfd; /* File descriptor for main journal */663 sqlite3_file *sjfd; /* File descriptor for sub-journal */664 i64 journalOff; /* Current write offset in the journal file */665 i64 journalHdr; /* Byte offset to previous journal header */666 sqlite3_backup *pBackup; /* Pointer to list of ongoing backup processes */667 PagerSavepoint *aSavepoint; /* Array of active savepoints */668 int nSavepoint; /* Number of elements in aSavepoint[] */669 u32 iDataVersion; /* Changes whenever database content changes */670 char dbFileVers[16]; /* Changes whenever database file changes */671 672 int nMmapOut; /* Number of mmap pages currently outstanding */673 sqlite3_int64 szMmap; /* Desired maximum mmap size */674 PgHdr *pMmapFreelist; /* List of free mmap page headers (pDirty) */675 /*676 ** End of the routinely-changing class members677 ***************************************************************************/678 679 u16 nExtra; /* Add this many bytes to each in-memory page */680 i16 nReserve; /* Number of unused bytes at end of each page */681 u32 vfsFlags; /* Flags for sqlite3_vfs.xOpen() */682 u32 sectorSize; /* Assumed sector size during rollback */683 Pgno mxPgno; /* Maximum allowed size of the database */684 Pgno lckPgno; /* Page number for the locking page */685 i64 pageSize; /* Number of bytes in a page */686 i64 journalSizeLimit; /* Size limit for persistent journal files */687 char *zFilename; /* Name of the database file */688 char *zJournal; /* Name of the journal file */689 int (*xBusyHandler)(void*); /* Function to call when busy */690 void *pBusyHandlerArg; /* Context argument for xBusyHandler */691 u32 aStat[4]; /* Total cache hits, misses, writes, spills */692#ifdef SQLITE_TEST693 int nRead; /* Database pages read */694#endif695 void (*xReiniter)(DbPage*); /* Call this routine when reloading pages */696 int (*xGet)(Pager*,Pgno,DbPage**,int); /* Routine to fetch a patch */697 char *pTmpSpace; /* Pager.pageSize bytes of space for tmp use */698 PCache *pPCache; /* Pointer to page cache object */699#ifndef SQLITE_OMIT_WAL700 Wal *pWal; /* Write-ahead log used by "journal_mode=wal" */701 char *zWal; /* File name for write-ahead log */702#endif703#ifdef SQLITE_ENABLE_SETLK_TIMEOUT704 sqlite3 *dbWal;705#endif706};707 708/*709** Indexes for use with Pager.aStat[]. The Pager.aStat[] array contains710** the values accessed by passing SQLITE_DBSTATUS_CACHE_HIT, CACHE_MISS711** or CACHE_WRITE to sqlite3_db_status().712*/713#define PAGER_STAT_HIT 0714#define PAGER_STAT_MISS 1715#define PAGER_STAT_WRITE 2716#define PAGER_STAT_SPILL 3717 718/*719** The following global variables hold counters used for720** testing purposes only. These variables do not exist in721** a non-testing build. These variables are not thread-safe.722*/723#ifdef SQLITE_TEST724int sqlite3_pager_readdb_count = 0; /* Number of full pages read from DB */725int sqlite3_pager_writedb_count = 0; /* Number of full pages written to DB */726int sqlite3_pager_writej_count = 0; /* Number of pages written to journal */727# define PAGER_INCR(v) v++728#else729# define PAGER_INCR(v)730#endif731 732 733 734/*735** Journal files begin with the following magic string. The data736** was obtained from /dev/random. It is used only as a sanity check.737**738** Since version 2.8.0, the journal format contains additional sanity739** checking information. If the power fails while the journal is being740** written, semi-random garbage data might appear in the journal741** file after power is restored. If an attempt is then made742** to roll the journal back, the database could be corrupted. The additional743** sanity checking data is an attempt to discover the garbage in the744** journal and ignore it.745**746** The sanity checking information for the new journal format consists747** of a 32-bit checksum on each page of data. The checksum covers both748** the page number and the pPager->pageSize bytes of data for the page.749** This cksum is initialized to a 32-bit random value that appears in the750** journal file right after the header. The random initializer is important,751** because garbage data that appears at the end of a journal is likely752** data that was once in other files that have now been deleted. If the753** garbage data came from an obsolete journal file, the checksums might754** be correct. But by initializing the checksum to random value which755** is different for every journal, we minimize that risk.756*/757static const unsigned char aJournalMagic[] = {758 0xd9, 0xd5, 0x05, 0xf9, 0x20, 0xa1, 0x63, 0xd7,759};760 761/*762** The size of the of each page record in the journal is given by763** the following macro.764*/765#define JOURNAL_PG_SZ(pPager) ((pPager->pageSize) + 8)766 767/*768** The journal header size for this pager. This is usually the same769** size as a single disk sector. See also setSectorSize().770*/771#define JOURNAL_HDR_SZ(pPager) (pPager->sectorSize)772 773/*774** The macro MEMDB is true if we are dealing with an in-memory database.775** We do this as a macro so that if the SQLITE_OMIT_MEMORYDB macro is set,776** the value of MEMDB will be a constant and the compiler will optimize777** out code that would never execute.778*/779#ifdef SQLITE_OMIT_MEMORYDB780# define MEMDB 0781#else782# define MEMDB pPager->memDb783#endif784 785/*786** The macro USEFETCH is true if we are allowed to use the xFetch and xUnfetch787** interfaces to access the database using memory-mapped I/O.788*/789#if SQLITE_MAX_MMAP_SIZE>0790# define USEFETCH(x) ((x)->bUseFetch)791#else792# define USEFETCH(x) 0793#endif794 795#ifdef SQLITE_DIRECT_OVERFLOW_READ796/*797** Return true if page pgno can be read directly from the database file798** by the b-tree layer. This is the case if:799**800** (1) the database file is open801** (2) the VFS for the database is able to do unaligned sub-page reads802** (3) there are no dirty pages in the cache, and803** (4) the desired page is not currently in the wal file.804*/805int sqlite3PagerDirectReadOk(Pager *pPager, Pgno pgno){806 assert( pPager!=0 );807 assert( pPager->fd!=0 );808 if( pPager->fd->pMethods==0 ) return 0; /* Case (1) */809 if( sqlite3PCacheIsDirty(pPager->pPCache) ) return 0; /* Failed (3) */810#ifndef SQLITE_OMIT_WAL811 if( pPager->pWal ){812 u32 iRead = 0;813 (void)sqlite3WalFindFrame(pPager->pWal, pgno, &iRead);814 if( iRead ) return 0; /* Case (4) */815 }816#else817 UNUSED_PARAMETER(pgno);818#endif819 assert( pPager->fd->pMethods->xDeviceCharacteristics!=0 );820 if( (pPager->fd->pMethods->xDeviceCharacteristics(pPager->fd)821 & SQLITE_IOCAP_SUBPAGE_READ)==0 ){822 return 0; /* Case (2) */823 }824 return 1;825}826#endif827 828#ifndef SQLITE_OMIT_WAL829# define pagerUseWal(x) ((x)->pWal!=0)830#else831# define pagerUseWal(x) 0832# define pagerRollbackWal(x) 0833# define pagerWalFrames(v,w,x,y) 0834# define pagerOpenWalIfPresent(z) SQLITE_OK835# define pagerBeginReadTransaction(z) SQLITE_OK836#endif837 838#ifndef NDEBUG839/*840** Usage:841**842** assert( assert_pager_state(pPager) );843**844** This function runs many asserts to try to find inconsistencies in845** the internal state of the Pager object.846*/847static int assert_pager_state(Pager *p){848 Pager *pPager = p;849 850 /* State must be valid. */851 assert( p->eState==PAGER_OPEN852 || p->eState==PAGER_READER853 || p->eState==PAGER_WRITER_LOCKED854 || p->eState==PAGER_WRITER_CACHEMOD855 || p->eState==PAGER_WRITER_DBMOD856 || p->eState==PAGER_WRITER_FINISHED857 || p->eState==PAGER_ERROR858 );859 860 /* Regardless of the current state, a temp-file connection always behaves861 ** as if it has an exclusive lock on the database file. It never updates862 ** the change-counter field, so the changeCountDone flag is always set.863 */864 assert( p->tempFile==0 || p->eLock==EXCLUSIVE_LOCK );865 assert( p->tempFile==0 || pPager->changeCountDone );866 867 /* If the useJournal flag is clear, the journal-mode must be "OFF".868 ** And if the journal-mode is "OFF", the journal file must not be open.869 */870 assert( p->journalMode==PAGER_JOURNALMODE_OFF || p->useJournal );871 assert( p->journalMode!=PAGER_JOURNALMODE_OFF || !isOpen(p->jfd) );872 873 /* Check that MEMDB implies noSync. And an in-memory journal. Since874 ** this means an in-memory pager performs no IO at all, it cannot encounter875 ** either SQLITE_IOERR or SQLITE_FULL during rollback or while finalizing876 ** a journal file. (although the in-memory journal implementation may877 ** return SQLITE_IOERR_NOMEM while the journal file is being written). It878 ** is therefore not possible for an in-memory pager to enter the ERROR879 ** state.880 */881 if( MEMDB ){882 assert( !isOpen(p->fd) );883 assert( p->noSync );884 assert( p->journalMode==PAGER_JOURNALMODE_OFF885 || p->journalMode==PAGER_JOURNALMODE_MEMORY886 );887 assert( p->eState!=PAGER_ERROR && p->eState!=PAGER_OPEN );888 assert( pagerUseWal(p)==0 );889 }890 891 /* If changeCountDone is set, a RESERVED lock or greater must be held892 ** on the file.893 */894 assert( pPager->changeCountDone==0 || pPager->eLock>=RESERVED_LOCK );895 assert( p->eLock!=PENDING_LOCK );896 897 switch( p->eState ){898 case PAGER_OPEN:899 assert( !MEMDB );900 assert( pPager->errCode==SQLITE_OK );901 assert( sqlite3PcacheRefCount(pPager->pPCache)==0 || pPager->tempFile );902 break;903 904 case PAGER_READER:905 assert( pPager->errCode==SQLITE_OK );906 assert( p->eLock!=UNKNOWN_LOCK );907 assert( p->eLock>=SHARED_LOCK );908 break;909 910 case PAGER_WRITER_LOCKED:911 assert( p->eLock!=UNKNOWN_LOCK );912 assert( pPager->errCode==SQLITE_OK );913 if( !pagerUseWal(pPager) ){914 assert( p->eLock>=RESERVED_LOCK );915 }916 assert( pPager->dbSize==pPager->dbOrigSize );917 assert( pPager->dbOrigSize==pPager->dbFileSize );918 assert( pPager->dbOrigSize==pPager->dbHintSize );919 assert( pPager->setSuper==0 );920 break;921 922 case PAGER_WRITER_CACHEMOD:923 assert( p->eLock!=UNKNOWN_LOCK );924 assert( pPager->errCode==SQLITE_OK );925 if( !pagerUseWal(pPager) ){926 /* It is possible that if journal_mode=wal here that neither the927 ** journal file nor the WAL file are open. This happens during928 ** a rollback transaction that switches from journal_mode=off929 ** to journal_mode=wal.930 */931 assert( p->eLock>=RESERVED_LOCK );932 assert( isOpen(p->jfd)933 || p->journalMode==PAGER_JOURNALMODE_OFF934 || p->journalMode==PAGER_JOURNALMODE_WAL935 );936 }937 assert( pPager->dbOrigSize==pPager->dbFileSize );938 assert( pPager->dbOrigSize==pPager->dbHintSize );939 break;940 941 case PAGER_WRITER_DBMOD:942 assert( p->eLock==EXCLUSIVE_LOCK );943 assert( pPager->errCode==SQLITE_OK );944 assert( !pagerUseWal(pPager) );945 assert( p->eLock>=EXCLUSIVE_LOCK );946 assert( isOpen(p->jfd)947 || p->journalMode==PAGER_JOURNALMODE_OFF948 || p->journalMode==PAGER_JOURNALMODE_WAL949 || (sqlite3OsDeviceCharacteristics(p->fd)&SQLITE_IOCAP_BATCH_ATOMIC)950 );951 assert( pPager->dbOrigSize<=pPager->dbHintSize );952 break;953 954 case PAGER_WRITER_FINISHED:955 assert( p->eLock==EXCLUSIVE_LOCK );956 assert( pPager->errCode==SQLITE_OK );957 assert( !pagerUseWal(pPager) );958 assert( isOpen(p->jfd)959 || p->journalMode==PAGER_JOURNALMODE_OFF960 || p->journalMode==PAGER_JOURNALMODE_WAL961 || (sqlite3OsDeviceCharacteristics(p->fd)&SQLITE_IOCAP_BATCH_ATOMIC)962 );963 break;964 965 case PAGER_ERROR:966 /* There must be at least one outstanding reference to the pager if967 ** in ERROR state. Otherwise the pager should have already dropped968 ** back to OPEN state.969 */970 assert( pPager->errCode!=SQLITE_OK );971 assert( sqlite3PcacheRefCount(pPager->pPCache)>0 || pPager->tempFile );972 break;973 }974 975 return 1;976}977#endif /* ifndef NDEBUG */978 979#ifdef SQLITE_DEBUG980/*981** Return a pointer to a human readable string in a static buffer982** containing the state of the Pager object passed as an argument. This983** is intended to be used within debuggers. For example, as an alternative984** to "print *pPager" in gdb:985**986** (gdb) printf "%s", print_pager_state(pPager)987**988** This routine has external linkage in order to suppress compiler warnings989** about an unused function. It is enclosed within SQLITE_DEBUG and so does990** not appear in normal builds.991*/992char *print_pager_state(Pager *p){993 static char zRet[1024];994 995 sqlite3_snprintf(1024, zRet,996 "Filename: %s\n"997 "State: %s errCode=%d\n"998 "Lock: %s\n"999 "Locking mode: locking_mode=%s\n"1000 "Journal mode: journal_mode=%s\n"1001 "Backing store: tempFile=%d memDb=%d useJournal=%d\n"1002 "Journal: journalOff=%lld journalHdr=%lld\n"1003 "Size: dbsize=%d dbOrigSize=%d dbFileSize=%d\n"1004 , p->zFilename1005 , p->eState==PAGER_OPEN ? "OPEN" :1006 p->eState==PAGER_READER ? "READER" :1007 p->eState==PAGER_WRITER_LOCKED ? "WRITER_LOCKED" :1008 p->eState==PAGER_WRITER_CACHEMOD ? "WRITER_CACHEMOD" :1009 p->eState==PAGER_WRITER_DBMOD ? "WRITER_DBMOD" :1010 p->eState==PAGER_WRITER_FINISHED ? "WRITER_FINISHED" :1011 p->eState==PAGER_ERROR ? "ERROR" : "?error?"1012 , (int)p->errCode1013 , p->eLock==NO_LOCK ? "NO_LOCK" :1014 p->eLock==RESERVED_LOCK ? "RESERVED" :1015 p->eLock==EXCLUSIVE_LOCK ? "EXCLUSIVE" :1016 p->eLock==SHARED_LOCK ? "SHARED" :1017 p->eLock==UNKNOWN_LOCK ? "UNKNOWN" : "?error?"1018 , p->exclusiveMode ? "exclusive" : "normal"1019 , p->journalMode==PAGER_JOURNALMODE_MEMORY ? "memory" :1020 p->journalMode==PAGER_JOURNALMODE_OFF ? "off" :1021 p->journalMode==PAGER_JOURNALMODE_DELETE ? "delete" :1022 p->journalMode==PAGER_JOURNALMODE_PERSIST ? "persist" :1023 p->journalMode==PAGER_JOURNALMODE_TRUNCATE ? "truncate" :1024 p->journalMode==PAGER_JOURNALMODE_WAL ? "wal" : "?error?"1025 , (int)p->tempFile, (int)p->memDb, (int)p->useJournal1026 , p->journalOff, p->journalHdr1027 , (int)p->dbSize, (int)p->dbOrigSize, (int)p->dbFileSize1028 );1029 1030 return zRet;1031}1032#endif1033 1034/* Forward references to the various page getters */1035static int getPageNormal(Pager*,Pgno,DbPage**,int);1036static int getPageError(Pager*,Pgno,DbPage**,int);1037#if SQLITE_MAX_MMAP_SIZE>01038static int getPageMMap(Pager*,Pgno,DbPage**,int);1039#endif1040 1041/*1042** Set the Pager.xGet method for the appropriate routine used to fetch1043** content from the pager.1044*/1045static void setGetterMethod(Pager *pPager){1046 if( pPager->errCode ){1047 pPager->xGet = getPageError;1048#if SQLITE_MAX_MMAP_SIZE>01049 }else if( USEFETCH(pPager) ){1050 pPager->xGet = getPageMMap;1051#endif /* SQLITE_MAX_MMAP_SIZE>0 */1052 }else{1053 pPager->xGet = getPageNormal;1054 }1055}1056 1057/*1058** Return true if it is necessary to write page *pPg into the sub-journal.1059** A page needs to be written into the sub-journal if there exists one1060** or more open savepoints for which:1061**1062** * The page-number is less than or equal to PagerSavepoint.nOrig, and1063** * The bit corresponding to the page-number is not set in1064** PagerSavepoint.pInSavepoint.1065*/1066static int subjRequiresPage(PgHdr *pPg){1067 Pager *pPager = pPg->pPager;1068 PagerSavepoint *p;1069 Pgno pgno = pPg->pgno;1070 int i;1071 for(i=0; i<pPager->nSavepoint; i++){1072 p = &pPager->aSavepoint[i];1073 if( p->nOrig>=pgno && 0==sqlite3BitvecTestNotNull(p->pInSavepoint, pgno) ){1074 for(i=i+1; i<pPager->nSavepoint; i++){1075 pPager->aSavepoint[i].bTruncateOnRelease = 0;1076 }1077 return 1;1078 }1079 }1080 return 0;1081}1082 1083#ifdef SQLITE_DEBUG1084/*1085** Return true if the page is already in the journal file.1086*/1087static int pageInJournal(Pager *pPager, PgHdr *pPg){1088 return sqlite3BitvecTest(pPager->pInJournal, pPg->pgno);1089}1090#endif1091 1092/*1093** Read a 32-bit integer from the given file descriptor. Store the integer1094** that is read in *pRes. Return SQLITE_OK if everything worked, or an1095** error code is something goes wrong.1096**1097** All values are stored on disk as big-endian.1098*/1099static int read32bits(sqlite3_file *fd, i64 offset, u32 *pRes){1100 unsigned char ac[4];1101 int rc = sqlite3OsRead(fd, ac, sizeof(ac), offset);1102 if( rc==SQLITE_OK ){1103 *pRes = sqlite3Get4byte(ac);1104 }1105 return rc;1106}1107 1108/*1109** Write a 32-bit integer into a string buffer in big-endian byte order.1110*/1111#define put32bits(A,B) sqlite3Put4byte((u8*)A,B)1112 1113 1114/*1115** Write a 32-bit integer into the given file descriptor. Return SQLITE_OK1116** on success or an error code is something goes wrong.1117*/1118static int write32bits(sqlite3_file *fd, i64 offset, u32 val){1119 char ac[4];1120 put32bits(ac, val);1121 return sqlite3OsWrite(fd, ac, 4, offset);1122}1123 1124/*1125** Unlock the database file to level eLock, which must be either NO_LOCK1126** or SHARED_LOCK. Regardless of whether or not the call to xUnlock()1127** succeeds, set the Pager.eLock variable to match the (attempted) new lock.1128**1129** Except, if Pager.eLock is set to UNKNOWN_LOCK when this function is1130** called, do not modify it. See the comment above the #define of1131** UNKNOWN_LOCK for an explanation of this.1132*/1133static int pagerUnlockDb(Pager *pPager, int eLock){1134 int rc = SQLITE_OK;1135 1136 assert( !pPager->exclusiveMode || pPager->eLock==eLock );1137 assert( eLock==NO_LOCK || eLock==SHARED_LOCK );1138 assert( eLock!=NO_LOCK || pagerUseWal(pPager)==0 );1139 if( isOpen(pPager->fd) ){1140 assert( pPager->eLock>=eLock );1141 rc = pPager->noLock ? SQLITE_OK : sqlite3OsUnlock(pPager->fd, eLock);1142 if( pPager->eLock!=UNKNOWN_LOCK ){1143 pPager->eLock = (u8)eLock;1144 }1145 IOTRACE(("UNLOCK %p %d\n", pPager, eLock))1146 }1147 pPager->changeCountDone = pPager->tempFile; /* ticket fb3b3024ea238d5c */1148 return rc;1149}1150 1151/*1152** Lock the database file to level eLock, which must be either SHARED_LOCK,1153** RESERVED_LOCK or EXCLUSIVE_LOCK. If the caller is successful, set the1154** Pager.eLock variable to the new locking state.1155**1156** Except, if Pager.eLock is set to UNKNOWN_LOCK when this function is1157** called, do not modify it unless the new locking state is EXCLUSIVE_LOCK.1158** See the comment above the #define of UNKNOWN_LOCK for an explanation1159** of this.1160*/1161static int pagerLockDb(Pager *pPager, int eLock){1162 int rc = SQLITE_OK;1163 1164 assert( eLock==SHARED_LOCK || eLock==RESERVED_LOCK || eLock==EXCLUSIVE_LOCK );1165 if( pPager->eLock<eLock || pPager->eLock==UNKNOWN_LOCK ){1166 rc = pPager->noLock ? SQLITE_OK : sqlite3OsLock(pPager->fd, eLock);1167 if( rc==SQLITE_OK && (pPager->eLock!=UNKNOWN_LOCK||eLock==EXCLUSIVE_LOCK) ){1168 pPager->eLock = (u8)eLock;1169 IOTRACE(("LOCK %p %d\n", pPager, eLock))1170 }1171 }1172 return rc;1173}1174 1175/*1176** This function determines whether or not the atomic-write or1177** atomic-batch-write optimizations can be used with this pager. The1178** atomic-write optimization can be used if:1179**1180** (a) the value returned by OsDeviceCharacteristics() indicates that1181** a database page may be written atomically, and1182** (b) the value returned by OsSectorSize() is less than or equal1183** to the page size.1184**1185** If it can be used, then the value returned is the size of the journal1186** file when it contains rollback data for exactly one page.1187**1188** The atomic-batch-write optimization can be used if OsDeviceCharacteristics()1189** returns a value with the SQLITE_IOCAP_BATCH_ATOMIC bit set. -1 is1190** returned in this case.1191**1192** If neither optimization can be used, 0 is returned.1193*/1194static int jrnlBufferSize(Pager *pPager){1195 assert( !MEMDB );1196 1197#if defined(SQLITE_ENABLE_ATOMIC_WRITE) \1198 || defined(SQLITE_ENABLE_BATCH_ATOMIC_WRITE)1199 int dc; /* Device characteristics */1200 