Replying to
@crankylinuxuser
@urbanfoxe @mattsheffield
No one do fine tune on context window. It's refinement on the probabilistic distribution with more tokens.
There is no training that is meaningful on such small scale. There no training that shows stability on self training. LLMs are just adjusting to the inputs.