Show HN: DynamicTune – Closed-form trajectory weight surgery from 4B into 0.8B
Show HN: DynamicTune – Closed-form trajectory weight surgery from 4B into 0.8B dsadawq3 / DynamicTune Public Notifications You must be signed in to change notification settings Fork 0 Star 0 Branches Tags Open more actions menu Latest commit History 6 Commits 6 Commits Folders and files Name Name Last commit message Last commit date benchmarks benchmarks configs configs data data docs docs faytuna_flow faytuna_flow scripts scripts tests tests .gitignore .gitignore README.md README.md pyproject.toml pyproject.toml uv.lock uv.lock Repository files navigation DynamicTune Challenging the trillion-token orthodoxy: cross-model hidden trajectory transport and closed-form weight surgery across architectures and model widths. Tested across radically different model families: modern hybrid Qwen3.5 (4B with d=2560 -> 0.8B with d=1024 ) and notoriously fragile GPT-2 (XL with d=1600 -> small with d=768 ).
This Research is relevant to the technology intelligence record because it involves Anthropic, Cohere, Perplexity, AMD. The source article should remain the factual reference for follow-up coverage.
- dsadawq3 / DynamicTune Public Notifications You must be signed in to change notification settings Fork 0 Star 0 Branches Tags Open more actions menu Latest commit History 6 Commits 6 Commits Folders and files Name Name Last commit message Last commit date benchmarks benchmarks configs configs data data docs docs faytuna_flow faytuna_flow scripts scripts tests tests .gitignore .gitignore README.md README.md pyproject.toml pyproject.toml uv.lock uv.lock Repository files navigation DynamicTune Challenging the trillion-token orthodoxy: cross-model hidden trajectory transport and closed-form weight surgery across architectures and model widths.
- Tested across radically different model families: modern hybrid Qwen3.5 (4B with d=2560 -> 0.8B with d=1024 ) and notoriously fragile GPT-2 (XL with d=1600 -> small with d=768 ).
- Breaking the Trillion-Token Orthodoxy The prevailing consensus in deep learning is that transferring capabilities from a larger teacher model to a smaller student demands billions or trillions of tokens, massive synthetic dataset pipelines, and weeks of GPU cluster compute running token-level cross-entropy or KL divergence minimization.
- DynamicTune challenges this dogma: A transformer stack is fundamentally a discrete dynamical system over depth: $h_{l+1} = h_l + f_l(h_l)$ .
- A capable teacher traces an informational velocity field through representation space.
- By aligning these trajectories through a local orthogonal Procrustes atlas and solving for closed-form weight updates in the student's MLP blocks, we can physically transfer teacher trajectory dynamics into the student without backpropagation or training runs.