This repository contains the BitStack-Llama-2-70B model as presented in BitStack: Fine-Grained Size Control for Compressed Large Language Models in Variable Memory Environments.
- Downloads last month
- 8
This model does not have enough activity to be deployed to Inference API (serverless) yet. Increase its social
visibility and check back later, or deploy to Inference Endpoints (dedicated)
instead.