extra chunking at 3072 characters? #4

Closed
opened 2026-09-15 13:51:37 -06:00 by tangent · 2 comments
Owner
  • 512 tokens is common
  • tokens are ~4 characters
  • longer context is a little better, especially if I am too verbose, which I am
  • 3072 should strike a balance of being small enough to make a meaningful difference while not so small as to drastically increase the amount of hashed/vectored stuff

I really should experiment with trying to accomplish tasks with the current setup before starting this though.

- 512 tokens is common - tokens are ~4 characters - longer context is a little better, especially if I am too verbose, which I am - 3072 should strike a balance of being small enough to make a meaningful difference while not so small as to drastically increase the amount of hashed/vectored stuff I really should experiment with trying to accomplish tasks with the current setup before starting this though.
Author
Owner

I think it's important to have multiple sizes combined, and since everything is working well with a max size, I'm implementing this by adding an optional target_chunk_size to generate these as well.

I think it's important to have multiple sizes combined, and since everything is working well with a max size, I'm implementing this by adding an optional `target_chunk_size` to generate these as well.
Author
Owner

There is no way for the script to recognize whether all target sizes have been met, so using this requires clearing the memory. I still haven't actually started with the goals yet though, so this is perfectly fine.

There is no way for the script to recognize whether all target sizes have been met, so using this requires clearing the memory. I still haven't actually *started* with the goals yet though, so this is perfectly fine.
Sign in to join this conversation.
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: tools/llm-tricks#4