~bzr-pqm/bzr/bzr.dev

« back to all changes in this revision

Viewing changes to bzrlib/revfile.py

Committer: Martin Pool
Date: 2005-05-11 01:09:41 UTC
Revision ID: mbp@sourcefrog.net-20050511010941-6438198e54086ddb

todo

files added:
bzrlib/atomicfile.py

bzrlib/help.py

bzrlib/log.py

bzrlib/statcache.py

bzrlib/status.py

contrib

contrib/add-bzr-to-baz

contrib/bash

contrib/bash/bzr

contrib/fortune

contrib/zsh

contrib/zsh/_bzr

doc/revfile-annotation.txt

doc/revfile.txt

doc/switch-in-branch.txt

files removed:
doc/faq.txt

doc/quickref.txt

test.sh

files modified:
.bzrignore

.rsyncexclude

NEWS

README

TODO

bzrlib/__init__.py

bzrlib/add.py

bzrlib/branch.py

bzrlib/commands.py

bzrlib/diff.py

bzrlib/errors.py

bzrlib/inventory.py

bzrlib/osutils.py

bzrlib/remotebranch.py

bzrlib/revfile.py

bzrlib/tests.py

bzrlib/textui.py

bzrlib/trace.py

doc/Makefile

doc/index.txt

doc/merge.txt

elementtree/ElementTree.py

testbzr

Show diffs side-by-side

added added

removed removed

bzrlib/revfile.py

is that sequence numbers are stable references. But not every

repository in the world will assign the same sequence numbers,

therefore the SHA-1 is the only universally unique reference.

This is meant to scale to hold 100,000 revisions of a single file, by

which time the index file will be ~4.8MB and a bit big to read

sequentially.

Some of the reserved fields could be used to implement a (semi?)

balanced tree indexed by SHA1 so we can much more efficiently find the

index associated with a particular hash. For 100,000 revs we would be

able to find it in about 17 random reads, which is not too bad.

This performs pretty well except when trying to calculate deltas of

really large files. For that the main thing would be to plug in

something faster than difflib, which is after all pure Python.

Another approach is to just store the gzipped full text of big files,

though perhaps that's too perverse?

The iter method here will generally read through the whole index file

in one go. With readahead in the kernel and python/libc (typically

128kB) this means that there should be no seeks and often only one

Older »