when a source is changed, its hash will be different, and this should be checked BEFORE doing any work
hash and path together should identify source documents, while pieces are stored by hash alone
store two indexes - a fast one with just hashes and vectors, and a full information one used for refresh and identifying sources
stored hashes are not checked for source deletion - the data store can only grow
when a source is moved, we can detect the change by checking against hashes to avoid duplicating work
when a source is moved and changed simultaneously, we cannot detect this, and it will be stored as new information; this is a flaw in the system that I do not wish to avoid right now
a wipe command should be allowed to regenerate everything from scratch and wipe old old data; this is the only way to remove source information and should only be used if LLMs keep pulling from outdated information created by a simultaneous move and modification of a source file
- [x] when a source is changed, its hash will be different, and this should be checked BEFORE doing any work
- [ ] hash and path together should identify source documents, while pieces are stored by hash alone
- [ ] store two indexes - a fast one with just hashes and vectors, and a full information one used for refresh and identifying sources
- [x] stored hashes are not checked for source deletion - the data store can only grow
- [ ] when a source is moved, we can detect the change by checking against hashes to avoid duplicating work
- [ ] when a source is moved and changed simultaneously, we cannot detect this, and it will be stored as new information; this is a flaw in the system that I do not wish to avoid right now
- [ ] a wipe command should be allowed to regenerate everything from scratch and wipe old old data; this is the only way to remove source information and should only be used if LLMs keep pulling from outdated information created by a simultaneous move and modification of a source file
I kind of ignored the stuff about trying to track file names and hashes, because having multiple file names aimed at the same hash isn't really bothersome to me right now. I might regret this later, but it shouldn't be too hard to clean up missing file names. That does almost go against the "never delete anything" but that's all content-addressed anyhow.
I kind of ignored the stuff about trying to track file names and hashes, because having multiple file names aimed at the same hash isn't really bothersome to me right now. I might regret this later, but it shouldn't be too hard to clean up missing file names. That does almost go against the "never delete anything" but that's all content-addressed anyhow.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
I kind of ignored the stuff about trying to track file names and hashes, because having multiple file names aimed at the same hash isn't really bothersome to me right now. I might regret this later, but it shouldn't be too hard to clean up missing file names. That does almost go against the "never delete anything" but that's all content-addressed anyhow.
Should be finished as of
c5aed984f8