BitaHub Model Deployment in Practice: Local Download, File System Upload, and GPU Mounting

A complete workflow for reducing model preparation costs on the BitaHub compute platform by leveraging its 200GB of free file storage: download models locally, upload via the command line, and mount them onto GPU instances.

BitaHubGPURTX 4090Hugging FaceModelScopehfmodel downloadmodel deploymentmodel inference

Overview

I previously used AutoDL quite a bit, and recently started using BitaHub (a compute platform offering A100 and RTX 4090 GPUs). The platform has relatively abundant RTX 4090 resources and slightly cheaper pricing.

For me, the main purpose of using this kind of compute platform is to deploy and test various open-source models, such as Qwen3.8-27B-Uncensored-GGUF, Qwen-Image-2.1, Comfy-Org/MiniMax-H3, and so on.

These scenarios all involve a necessary step: model download. Model weight files are typically tens of GB in size, for example:

If every time you start a GPU instance, download the model, wait for tens of GB to finish downloading, and only then begin inference, the time spent beforehand is purely downloading, which uses none of the GPU's compute power β€” yet the download process still incurs GPU charges. That's somewhat uneconomical.

One nice thing about AutoDL is that when you don't need GPU scenarios, you can shut down the instance and use no-GPU mode. No-GPU mode uses 0.5 core, 2GB memory, and no GPU card, at a flat price of Β₯0.1/hour, with no impact on the instance's data before or after. See: AutoDL money-saving tips.

BitaHub doesn't offer this feature, but it does provide a separate file storage feature. You can use it to download models locally, upload them to BitaHub's file storage, and then mount models from the file storage β€” this avoids incurring GPU charges while downloading models.

The core principle is simple:

Anything that can be done before the GPU instance starts shouldn't consume GPU time.

Compute and Storage

BitaHub provides different types of compute resources including RTX 4090 24GB, NVIDIA A100 80GB, and CPU, along with single-GPU and multi-GPU configurations.

I currently mainly use the RTX 4090 24GB. Taking the information shown on the platform at the time I use it as an example:

ResourceSpec / Price
RTX 409024GB VRAM
4090 singleΒ₯1.68 / hour
Instance RAM56GB
GPU config1 / 2 / 4 / 8 GPUs
A10080GB VRAM
A100 singleΒ₯6.38 / hour

These are the resources and prices currently displayed in real time on the platform. For actual use, refer to the information shown on the instance creation page.

The free system disk capacity provided for the dev machine is 30GB, and it can only be set up to a maximum of 100GB. The additional 70GB of system disk costs Β₯0.07/day/GB, totaling Β₯4.90/day, regardless of whether the machine is shut down.

Therefore, it's best to store models using the platform-provided file storage. Each user has a 200GB free storage quota, and the excess is charged at Β₯0.01/GB/day. For individual users, 200GB can already hold multiple sets of models simultaneously.

Plain text
Qwen-Image-2.1                         β‰ˆ 33 GB
MiniMax-H3 related quantized models & components β‰ˆ 70 GB
Qwen3.8-27B Uncensored GGUF Q5         β‰ˆ 22 GB
------------------------------------------------
Total                                   β‰ˆ 125 GB

The remaining space can also hold VAE, Text Encoder, LoRA, or other files. Storing large model weights in a separate file system and mounting them onto GPU instances on demand is an appropriate way to use it.

Model Preparation Workflow

After actual use, the workflow settles into the following:

Plain text
1. Confirm the model to deploy
        ↓
2. Check ModelScope first (very fast access in China)
        ↓
3. If not on ModelScope, check Hugging Face (some Uncensored models are usually not available on ModelScope)
        ↓
4. Download to a local directory
        ↓
5. Create a file system on BitaHub (first time)
        ↓
6. Create a model on BitaHub; a file system must be created first
        ↓
7. Upload to the specified model using the bita CLI
        ↓
8. Create a GPU development environment
        ↓
9. Mount the file system / model
        ↓
10. Start inference

This way, before GPU billing actually begins, the model files β€” tens or even hundreds of GB β€” are already fully prepared.

Note: if an account hasn't logged in for over 6 months and has a zero balance, or if an account is in arrears for over 15 days, the platform will clean up the account data.

Tools

You mainly need three command-line tools locally:

  • hf: download models from the Hugging Face Hub;

  • modelscope: download models from the ModelScope (MoDa) community;

  • bita: a command-line fast-upload tool for uploading local models to BitaHub.

Hugging Face CLI

The huggingface_hub Python package ships with a built-in command-line tool: hf. This tool allows you to interact directly with the Hugging Face Hub from the terminal, such as: logging in, creating repositories, uploading and downloading, etc.

Installation

  • macOS and Linux
Bash
curl -LsSf https://hf.co/cli/install.sh | bash
  • Windows
Bash
powershell -ExecutionPolicy ByPass -c "irm https://hf.co/cli/install.ps1 | iex"

The installer also installs the global hf-cli skill for use by any agent that reads the ~/.agents/skills directory. Use the --exclude-skill flag to avoid installing it:

Bash
# macOS and Linux
curl -LsSf https://hf.co/cli/install.sh | bash -s -- --exclude-skill
 
# Windows
powershell -ExecutionPolicy ByPass -c "& ([scriptblock]::Create((irm https://hf.co/cli/install.ps1))) -ExcludeSkill"

Usage

Basic download command structure:

Bash
hf download <repo-id> [files...] --local-dir <directory>
  • Download the entire repository to a directory:
Bash
hf download Qwen/Qwen-Image-2.1 --local-dir ~/MyFiles/models/Qwen-Image-2.1
  • Download a single file to a directory:
Bash
hf download JonathanColetti/Qwen3.8-27B-Uncensored-GGUF Qwen3.8-27B-Uncensored-Q5_K_M.gguf --local-dir ~/MyFiles/models/Qwen3.8-27B-Uncensored-GGUF
  • Download multiple specified files at once. This is common with GGUF repositories, which provide many quantized versions such as IQ2, Q4, Q5, Q6, Q8, etc. In practice you only need one main model matching your GPU VRAM plus auxiliary files, not the entire repository:
Bash
hf download JonathanColetti/Qwen3.8-27B-Uncensored-GGUF \
  Qwen3.8-27B-Uncensored-Q5_K_M.gguf \
  Qwen3.8-27B-Uncensored-draft-Q4_0.gguf \
  Qwen3.8-27B-Uncensored-draft-Q8_0.gguf \
  mmproj-Qwen3.8-27B-Uncensored-F16.gguf \
  --local-dir ~/MyFiles/models/Qwen3.8-27B-Uncensored-GGUF
  • Use --include to select the files you need according to a glob pattern:
Bash
hf download <repo-id> --include "*.json" "*.safetensors" --local-dir ./models/model-name/
  • Use --exclude when you need most of the files but want to exclude certain ones. For example, if a repository provides both .safetensors and .bin weights and you only need safetensors, you can directly exclude .bin:
Bash
hf download <repo-id> --exclude "*.bin" --local-dir ./models/model-name/

If you plan to upload to BitaHub, you must use --local-dir. If you simply run:

Bash
hf download <repo-id>

Hugging Face uses its own Hub Cache structure by default, and files are usually located at:

Bash
~/.cache/huggingface/hub/

The structure uses symlinks pointing to the actual files. This is Hugging Face's default cache structure, but if your goal is to upload to a remote server, it isn't well-suited for file organization.

Bash
tree -l models--mobiuslabsgmbh--faster-whisper-large-v3-turbo
models--mobiuslabsgmbh--faster-whisper-large-v3-turbo
β”œβ”€β”€ blobs
β”‚   β”œβ”€β”€ 0351d1d6870005e865747b781b5d7c23ea0459cd
β”‚   β”œβ”€β”€ 0adcd01e7c237205d593b707e66dd5d7bc785d2d
β”‚   β”œβ”€β”€ 17456db595adc78a973f97d69d8cb50bc87c0b1c
β”‚   β”œβ”€β”€ 931c77a740890c46365c7ae0c9d350ba3cca908f
β”‚   └── e76620f83d5f5b69efd3d87e3dc180c1bd21df9fbebacfd4335e5e1efcc018da
β”œβ”€β”€ refs
β”‚   └── main
└── snapshots
    └── 0a363e9161cbc7ed1431c9597a8ceaf0c4f78fcf
        β”œβ”€β”€ config.json -> ../../blobs/0351d1d6870005e865747b781b5d7c23ea0459cd
        β”œβ”€β”€ model.bin -> ../../blobs/e76620f83d5f5b69efd3d87e3dc180c1bd21df9fbebacfd4335e5e1efcc018da
        β”œβ”€β”€ preprocessor_config.json -> ../../blobs/931c77a740890c46365c7ae0c9d350ba3cca908f
        β”œβ”€β”€ tokenizer.json -> ../../blobs/17456db595adc78a973f97d69d8cb50bc87c0b1c
        └── vocabulary.json -> ../../blobs/0adcd01e7c237205d593b707e66dd5d7bc785d2d
 
5 directories, 11 files

It's best to specify directly:

Bash
# ~/models/<model-name>/ is the local path
--local-dir ~/models/<model-name>/

The resulting directory structure is basically the same as the model repository itself, making it easier to inspect, package, upload, and later load directly.

Note that currently, even when using --local-dir, Hugging Face will create a .cache/huggingface/ metadata directory in the target directory, used to determine whether files need to be re-downloaded.

After the model download completes and you've confirmed you no longer need to sync updates, this metadata directory can be deleted.

ModelScope CLI

You need to use the ModelScope SDK to download models. Official documentation

Installation

Bash
pip install -U modelscope
# Confirm the installation succeeded
modelscope --version

Usage

ModelScope's logic is very similar to Hugging Face's. The basic command-line download format is:

Bash
modelscope download --model <model-id> --local_dir <directory>

The parameter naming is slightly different between the two:

Plain text
Hugging Face    --local-dir
ModelScope      --local_dir

One uses -, the other uses _.

  • Default download location: ~/.cache/modelscope/hub. To change it, set the environment variable MODELSCOPE_CACHE. On macOS, add the following to ~/.zshrc:
Bash
export MODELSCOPE_CACHE=/path/to/your/cache
  • <model_id> is the model ID, obtained from the model's homepage, consisting of the organization name and model name, e.g. Qwen/Qwen-Image-2.1;

  • Downloading private models requires login;

  • Download the entire repository to a directory:

Bash
modelscope download --model Qwen/Qwen-Image-2.1 --local_dir /path/to/your/models/
  • Download a single file to a directory:
Bash
modelscope download --model Qwen/Qwen-Image-2.1 README.md --local_dir /path/to/your/models/
  • To download multiple specified files, write the file names one after another:
Bash
modelscope download \
  --model <model-id> \
  README.md \
  model_index.json \
  --local_dir /path/to/your/models/
  • Use --include to select the files you need according to a glob pattern:
Bash
modelscope download \
  --model <model-id> \
  --include "*.safetensors" \
  --local_dir /path/to/your/models/

You can also specify multiple patterns:

Bash
modelscope download \
  --model <model-id> \
  --include "*.json" "*.safetensors" \
  --local_dir ./models/model-name/
  • Use --exclude when you need most of the files but want to exclude certain ones. For example, if a repository provides both .safetensors and .bin weights and you only need safetensors, you can directly exclude .bin:
Bash
modelscope download \
  --model <model-id> \
  --exclude "*.bin" \
  --local_dir /path/to/your/models/
  • Parameters:
ParameterShortTypeDefaultDescription
repo_id-str-Positional arg, repository ID (optional, can also be set via --model)
files-str-Positional arg, specifies the files to download (supports multiple)
--model-strNoneModel ID (mutually exclusive with --dataset)
--dataset-strNoneDataset ID (mutually exclusive with --model)
--repo-type-choicemodelRepository type (model/dataset), used with the positional repo_id
--revision-strNoneVersion / branch / tag
--cache_dir-strNoneCache directory
--local_dir-strNoneLocal directory (takes precedence over cache_dir)
--include-listNoneFile glob patterns to include
--exclude-listNoneFile glob patterns to exclude
--token-strNoneAccess token (required for private models)
--endpoint-strNoneModelScope service endpoint
--max-workers-intdefaultMaximum number of concurrent download threads

BitaHub CLI

BitaHub provides bita client versions for Windows, Linux, macOS Intel, and macOS Apple Silicon. I use macOS Apple Silicon. After installation, confirm:

Bash
bita --help

BitaHub officially recommends that on macOS you ensure /usr/local/bin/ exists before installation, then install the .pkg for your architecture.

Bash
mkdir -p /usr/local/bin/

Once the model is prepared, you can upload it to BitaHub. First, you need to create a file system in BitaHub's File Storage.

For models of tens of GB, using browser upload directly is not recommended. BitaHub also officially recommends using the bita command-line fast-upload tool when a single file exceeds 10GB, when there are many files, or when you need to upload from a local terminal.

  1. Log in:
Bash
bita login -e https://www.bitahub.com

The CLI uses browser-based authorization. The terminal displays an authorization URL and a one-time authorization code. Complete the authorization in the browser β€” you don't need to enter your account password directly in the terminal.

  1. Get the upload command:

Go to BitaHub's Storage Management page, select the corresponding file system, and click the Upload button. The page will automatically generate something like:

Bash
bita upload \
  -e https://www.bitahub.com \
  -b <bucket-id> \
  -o / \
  -l <local file or folder path>

Where:

Plain text
-e    BitaHub platform address
-b    bucket id corresponding to the current file system
-o    target directory in the file system for upload
-l    local file or folder path

In practice, you don't need to manually enter the bucket-id β€” it's safest to directly copy the command generated by BitaHub's current upload page.

Mount and Run the Model

First, complete all the preparation work:

Plain text
Local model download complete
        ↓
File system created
        ↓
Model uploaded
        ↓
BitaHub model created

After everything is prepared, then create the GPU development environment.

The model deployment workflow officially provided by BitaHub is also: upload models to the file system first, create models from these files, and finally mount the corresponding model when creating the development environment.

When creating the development environment, select the previously created file system or model under Storage Mount. After the model is mounted, you can read the model files directly from the mount directory inside the container, for example:

Bash
/oceanz/models/Qwen/Qwen-Image-2.1

The actual path depends on the name used when creating the model and the mount configuration. Afterwards, the code can load using the local path directly, for example:

Python
pipe = QwenImage21Pipeline.from_pretrained(
    "/oceanz/models/Qwen/Qwen-Image-2.1",
    torch_dtype=torch.bfloat16,
    local_files_only=True
)

For GGUF with llama.cpp:

Bash
llama cli -m /oceanz/models/Qwen/Qwen3.8-27B-Uncensored/Qwen3.8-27B-Uncensored-Q5_K_M.gguf

If it's a vision model, also mount the corresponding multimodal projector:

Bash
llama cli -m /oceanz/models/Qwen/Qwen3.8-27B-Uncensored/Qwen3.8-27B-Uncensored-Q5_K_M.gguf  --mmproj /oceanz/models/Qwen/Qwen3.8-27B-Uncensored/mmproj-Qwen3.8-27B-Uncensored-F16.gguf

This way, once the GPU instance starts, you can basically begin environment setup and inference directly, without waiting for tens of GB of model downloads.

No table of contents