write itemscompare items
. . .
addr datalenmem-idaddr datalenmem-id
. . .
addr datalenmem-idaddr datalenmem-id
M i n i t r a n s a c t i o n
read items
. . .
addr lenmem-idaddr lenmem-idclass Minitransaction {public:void cmp(memid,addr,len,data); // addvoid write(memid,addr,len,data); // addint exec_and_commit(); // execute and commit};
cmp itemwrite item
void read(memid,addr,len,buf); // add
read item
...t = new Minitransaction;t->cmp(memid, addr, len, data);t->write(memid, addr, len, newdata);status = t->exec_and_commit();...
API Example
··
check data indicated by(equality comparison)if all match thenretrieve data indicated bymodify data indicated by
compare itemsread itemswrite items
Semantics of a minitransaction
Figure 2: Minitransactions have compare items, read items, and write items. Compare items are locations to compare against givenvalues, while read items are locations to read and write items arelocations to update, if all comparisons match. All items are speci- fied before the minitransaction starts executing. The example codecreates a minitransaction with one compare and one write item onthe same location—a compare-and-swap operation. Methods
cmp
,
read
, and
write
populate a minitransaction without communicationwith memory nodes until
exec_and_commit
is invoked.
and commit. Roughly speaking, a coordinator executes a transac-tion by asking participants to perform one or more transaction
ac-tions
, such as retrieving or modifying data items. At the end of the transaction, the coordinator executes two-phase commit. In thefirst phase, the coordinator asks all participants if they are ready tocommit. If they all vote yes, in the second phase the coordinatortells them to commit; otherwise the coordinator tells them to abort.In Sinfonia, coordinators are application nodes and participants arememory nodes.We observe that it is possible to optimize the execution of sometransactions, as follows. If the transaction’s last action does notaffect the coordinator’s decision to abort or commit then the coor-dinator can piggyback this last action onto the first phase of two-phase commit (e.g., this is the case if this action is a data update).This optimization does not affect the transaction semantics andsaves a communication round-trip.Even if the transaction’s last action affects the coordinator’s de-cision to abort or commit, if the participant knows how the coor-dinator makes this decision, then we can also piggyback the actiononto the commit protocol. For example, if the last action is a readand the participant knows that the coordinator will abort if the readreturns zero (and will commit otherwise), then the coordinator canpiggyback this action onto two-phase commit and the participantcan read the item and adjust its vote to abort if the result is zero.In fact, it might be possible to piggyback the entire transactionexecution onto the commit protocol. We designed minitransactionsso that this is always the case and found that it is still possible toget fairly powerful transactions.More precisely, a minitransaction (Figure 2) consists of a set of
compare items
, a set of
read items
, and a set of
write items
. Eachitem specifies a memory node and an address range within thatmemory node; compare and write items also include data. Itemsare chosen before the minitransaction starts executing. Upon exe-cution, a minitransaction does the following: (1) compare the loca-tions in the compare items, if any, against the data in the compareitems (equality comparison), (2) if all comparisons succeed, or if there are no compare items, return the locations in the read itemsand write to the locations in the write items, and (3) if some com-parison fails, abort. Thus, the compare items control whether theminitransaction commits or aborts, while the read and write itemsdetermine what data the minitransaction returns and updates.Minitransactions are a powerful primitive for handling dis-tributed data. Examples of minitransactions include the following:1.
Swap.
A read item returns the old value and a write itemreplaces it.2.
Compare-and-swap.
A compare item compares the currentvalue against a constant; if equal, a write item replaces it.3.
Atomic read of many data.
Done with multiple read items.4.
Acquire a lease.
A compare item checks if a location is setto 0; if so, a write item sets it to the (non-zero)
id
of theleaseholder and another write item sets the time of lease.5.
Acquire multiple leases atomically.
Same as above, exceptthat there are multiple compare items and write items. Notethat each lease can be in a different memory node.6.
Change data if lease is held.
A compare item checks that alease is held and, if so, write items update data.A frequent minitransaction idiom is to use compare items to vali-date data and, ifdata is valid, use writeitems to apply some changesto the same or different data. These minitransactions are commonin SinfoniaFS: the file system caches inodes and metadata aggres-sively at application nodes, and relevant cached entries are vali-dated before modifying the file system. For example, writing to afile requires validating a cached copy of the file’s inode and chain-ing list (the list of blocks comprising the file) and, if they are valid,modifying the appropriate file block. This is done with compareitems and write items in a minitransaction. Figure 7 shows a mini-transaction used by SinfoniaFS to set a file’s attributes.Another minitransaction idiom is to have only compare items tovalidate data, without read or write items. Such a minitransactionmodifies no data, regardless of whether it commits or aborts. Butif it commits, the application node knows that all comparisons suc-ceeded and so the validations were successful. SinfoniaFS usesthis type of minitransaction to validate cached data for read-onlyfile system operations, such as stat (NFS’s getattr).In Section 4 we explain how minitransactions are executed andcommitted efficiently. It is worth noting that minitransactions canbe extended to include more general read-modify-write items (not just write items) and generic conditional items (not just compareitems) provided that each item can be executed at a single memorynode. For example, there could be an increment item that atomi-cally increments a location; and a minitransaction could have mul-tiple increment items, possibly at different memory nodes, to in-crement all of them together. These extensions were not neededfor the applications in this paper, but they may be useful for otherapplications.
3.4 Caching and consistency
Sinfonia does not cache data at application nodes, but providessupport for applications to do their own caching. Application-controlled caching has three clear advantages: First, there is greaterflexibilityon policies ofwhat tocache and what to evict. Second, asa result, cache utilization potentially improves, since applicationsknow their data better than what Sinfonia can infer. And third,Sinfonia becomes a simpler service to use because data accessedthrough Sinfonia is always current (not stale). Managing caches in
Leave a Comment