Implementation of the COLA framework proposed in the paper "Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization".
COLA is a three-stage framework for curating high-quality calibration data to preserve LLM capabilities during compression:
- Dataset Selection (Domain Correspondence): Selects datasets that align with the target deployment domain.
- Dataset Processing (Compositional Properties): Optimizes the compositional properties of selected datasets (sequence length, format, etc.).
- Sample Selection (Representativeness and Diversity in Activation Space): Selects samples that maximize representativeness and diversity in the model's activation space.
# Clone the repository
git clone https://anonymous.4open.science/r/COLA-7D2C
cd COLA
# Install the package
pip install -e .from transformers import AutoModelForCausalLM, AutoTokenizer
from cola import COLA
# Load LLM model and tokenizer
model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3-8b")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3-8b")
# Initialize COLA framework
cola = COLA(
model=model,
tokenizer=tokenizer,
available_datasets=["wikitext", "c4", "slimpajama-200k"],
target_capabilities=["commonsense", "math", "code"],
output_dir="./cola_output"
)
# Run COLA to generate calibration data
calibration_samples = cola.run(
num_samples=128, # Number of samples to select
sequence_length=2048, # Target sequence length
)You can also use the provided example script:
python run_cola.py \
--model_name_or_path meta-llama/Llama-3-8b \
--output_dir ./cola_output \
--num_samples 128 \
--sequence_length 2048 \
--target_capabilities commonsense math code \
--datasets wikitext c4 slimpajama-200k \
--deployment_type generalFor targeted deployment (focusing on a specific capability):
python run_cola.py \
--model_name_or_path meta-llama/Llama-3-8b \
--output_dir ./cola_output \
--num_samples 128 \
--sequence_length 2048 \
--target_capabilities commonsense math code \
--datasets wikitext c4 slimpajama-200k \
--deployment_type targeted \
--targeted_capability mathThis stage focuses on selecting datasets that align with the target deployment domain:
- Analyzes whether the compressed model is intended for general-purpose use or specialized tasks
- Selects a balanced mix of pre-training datasets for general-purpose deployment
- Prioritizes domain-matched datasets for targeted deployment
- Focuses on language alignment, subject coverage, and reasoning difficulty
This stage optimizes the compositional properties of the selected datasets:
- Optimizes sequence length (typically 2048 tokens for most methods)
- Enhances format by converting to Q&A format with explicit reasoning chains
- Filters out low-quality samples that could negatively impact compression
This stage selects individual samples to maximize representativeness and diversity in activation space:
- Extracts layer-wise activations from the uncompressed model
- Applies dimensionality reduction using random projection
- Clusters samples in the activation space using k-means
- Selects representative samples from each cluster
COLA is designed to be compatible with various post-training compression methods:
- SparseGPT
- Wanda
- LLM-Pruner
- RIA
- GPTQ
- AWQ
- SmoothQuant
- FlatQuant
To use COLA with these methods, simply generate the calibration data using COLA, then use it as input for your chosen compression method.
The calibration samples produced by COLA are saved in JSON format:
[
{
"text": "Question: How does photosynthesis work?\n\nReasoning:\nPhotosynthesis is the process used by plants, algae and certain bacteria to convert light energy, usually from the sun, into chemical energy in the form of glucose or other sugars. These are synthesized from carbon dioxide and water.\n\nThe process occurs in multiple steps:\n1. Light energy is absorbed by chlorophyll in the chloroplasts\n2. This energy is used to split water molecules, releasing oxygen\n3. The hydrogen from water and carbon dioxide from the air are used to form glucose\n4. Oxygen is released as a byproduct\n\nThe overall equation is:\n6CO₂ + 6H₂O + light energy → C₆H₁₂O₆ + 6O₂\n\nAnswer: Photosynthesis is the process where plants convert sunlight, water, and carbon dioxide into glucose and oxygen. Chlorophyll captures light energy, which powers chemical reactions that split water and reduce carbon dioxide to create sugar molecules, releasing oxygen as a byproduct.",
"dataset_name": "c4",
"capability_scores": {
"commonsense": 0.85,
"math": 0.42,
"code": 0.31
},
"format_enhanced": true,
"selection_index": 14,
"selection_method": "activation_clustering"
},
...
]Based on the paper's findings, optimal calibration data for capability preservation should:
- Have representative activation patterns for the target domain
- Maintain diversity in activation space to cover the model's full capabilities
- Include explicit reasoning chains for preserving reasoning capabilities
- Have appropriate sequence length (typically 2048 tokens for most methods)
- Use domain-matched data for targeted deployment scenarios
- Include mixed difficulty levels for a balance of specialized and general performance
MIT License