Overview
I previously used AutoDL quite a bit, and recently started using BitaHub (a compute platform offering A100 and RTX 4090 GPUs). The platform has relatively abundant RTX 4090 resources and slightly cheaper pricing.
For me, the main purpose of using this kind of compute platform is to deploy and test various open-source models, such as Qwen3.8-27B-Uncensored-GGUF, Qwen-Image-2.1, Comfy-Org/MiniMax-H3, and so on.
These scenarios all involve a necessary step: model download. Model weight files are typically tens of GB in size, for example:
- Qwen-Image-2.1 : about 33 GB;
- Comfy-Org/MiniMax-H3 and related quantized models: the entire workflow may require about 70 GB;
- Qwen3.8-27B Uncensored GGUF Q5 plus mmproj, draft model, and other files : about 20+ GB.
If every time you start a GPU instance, download the model, wait for tens of GB to finish downloading, and only then begin inference, the time spent beforehand is purely downloading, which uses none of the GPU's compute power β yet the download process still incurs GPU charges. That's somewhat uneconomical.
One nice thing about AutoDL is that when you don't need GPU scenarios, you can shut down the instance and use no-GPU mode. No-GPU mode uses 0.5 core, 2GB memory, and no GPU card, at a flat price of Β₯0.1/hour, with no impact on the instance's data before or after. See: AutoDL money-saving tips.
BitaHub doesn't offer this feature, but it does provide a separate file storage feature. You can use it to download models locally, upload them to BitaHub's file storage, and then mount models from the file storage β this avoids incurring GPU charges while downloading models.
The core principle is simple:
Anything that can be done before the GPU instance starts shouldn't consume GPU time.
Compute and Storage
BitaHub provides different types of compute resources including RTX 4090 24GB, NVIDIA A100 80GB, and CPU, along with single-GPU and multi-GPU configurations.
I currently mainly use the RTX 4090 24GB. Taking the information shown on the platform at the time I use it as an example:
| Resource | Spec / Price |
|---|---|
| RTX 4090 | 24GB VRAM |
| 4090 single | Β₯1.68 / hour |
| Instance RAM | 56GB |
| GPU config | 1 / 2 / 4 / 8 GPUs |
| A100 | 80GB VRAM |
| A100 single | Β₯6.38 / hour |
These are the resources and prices currently displayed in real time on the platform. For actual use, refer to the information shown on the instance creation page.
The free system disk capacity provided for the dev machine is 30GB, and it can only be set up to a maximum of 100GB. The additional 70GB of system disk costs Β₯0.07/day/GB, totaling Β₯4.90/day, regardless of whether the machine is shut down.
Therefore, it's best to store models using the platform-provided file storage. Each user has a 200GB free storage quota, and the excess is charged at Β₯0.01/GB/day. For individual users, 200GB can already hold multiple sets of models simultaneously.
Qwen-Image-2.1 β 33 GB
MiniMax-H3 related quantized models & components β 70 GB
Qwen3.8-27B Uncensored GGUF Q5 β 22 GB
------------------------------------------------
Total β 125 GBThe remaining space can also hold VAE, Text Encoder, LoRA, or other files. Storing large model weights in a separate file system and mounting them onto GPU instances on demand is an appropriate way to use it.
Model Preparation Workflow
After actual use, the workflow settles into the following:
1. Confirm the model to deploy
β
2. Check ModelScope first (very fast access in China)
β
3. If not on ModelScope, check Hugging Face (some Uncensored models are usually not available on ModelScope)
β
4. Download to a local directory
β
5. Create a file system on BitaHub (first time)
β
6. Create a model on BitaHub; a file system must be created first
β
7. Upload to the specified model using the bita CLI
β
8. Create a GPU development environment
β
9. Mount the file system / model
β
10. Start inferenceThis way, before GPU billing actually begins, the model files β tens or even hundreds of GB β are already fully prepared.
Note: if an account hasn't logged in for over 6 months and has a zero balance, or if an account is in arrears for over 15 days, the platform will clean up the account data.Tools
You mainly need three command-line tools locally:
-
hf: download models from the Hugging Face Hub; -
modelscope: download models from the ModelScope (MoDa) community; -
bita: a command-line fast-upload tool for uploading local models to BitaHub.
Hugging Face CLI
The huggingface_hub Python package ships with a built-in command-line tool: hf. This tool allows you to interact directly with the Hugging Face Hub from the terminal, such as: logging in, creating repositories, uploading and downloading, etc.
Installation
- macOS and Linux
curl -LsSf https://hf.co/cli/install.sh | bash- Windows
powershell -ExecutionPolicy ByPass -c "irm https://hf.co/cli/install.ps1 | iex"The installer also installs the global hf-cli skill for use by any agent that reads the ~/.agents/skills directory. Use the --exclude-skill flag to avoid installing it:
# macOS and Linux
curl -LsSf https://hf.co/cli/install.sh | bash -s -- --exclude-skill
# Windows
powershell -ExecutionPolicy ByPass -c "& ([scriptblock]::Create((irm https://hf.co/cli/install.ps1))) -ExcludeSkill"Usage
Basic download command structure:
hf download <repo-id> [files...] --local-dir <directory>- Download the entire repository to a directory:
hf download Qwen/Qwen-Image-2.1 --local-dir ~/MyFiles/models/Qwen-Image-2.1- Download a single file to a directory:
hf download JonathanColetti/Qwen3.8-27B-Uncensored-GGUF Qwen3.8-27B-Uncensored-Q5_K_M.gguf --local-dir ~/MyFiles/models/Qwen3.8-27B-Uncensored-GGUF- Download multiple specified files at once. This is common with GGUF repositories, which provide many quantized versions such as IQ2, Q4, Q5, Q6, Q8, etc. In practice you only need one main model matching your GPU VRAM plus auxiliary files, not the entire repository:
hf download JonathanColetti/Qwen3.8-27B-Uncensored-GGUF \
Qwen3.8-27B-Uncensored-Q5_K_M.gguf \
Qwen3.8-27B-Uncensored-draft-Q4_0.gguf \
Qwen3.8-27B-Uncensored-draft-Q8_0.gguf \
mmproj-Qwen3.8-27B-Uncensored-F16.gguf \
--local-dir ~/MyFiles/models/Qwen3.8-27B-Uncensored-GGUF- Use
--includeto select the files you need according to a glob pattern:
hf download <repo-id> --include "*.json" "*.safetensors" --local-dir ./models/model-name/- Use
--excludewhen you need most of the files but want to exclude certain ones. For example, if a repository provides both.safetensorsand.binweights and you only need safetensors, you can directly exclude.bin:
hf download <repo-id> --exclude "*.bin" --local-dir ./models/model-name/If you plan to upload to BitaHub, you must use --local-dir. If you simply run:
hf download <repo-id>Hugging Face uses its own Hub Cache structure by default, and files are usually located at:
~/.cache/huggingface/hub/The structure uses symlinks pointing to the actual files. This is Hugging Face's default cache structure, but if your goal is to upload to a remote server, it isn't well-suited for file organization.
tree -l models--mobiuslabsgmbh--faster-whisper-large-v3-turbo
models--mobiuslabsgmbh--faster-whisper-large-v3-turbo
βββ blobs
β βββ 0351d1d6870005e865747b781b5d7c23ea0459cd
β βββ 0adcd01e7c237205d593b707e66dd5d7bc785d2d
β βββ 17456db595adc78a973f97d69d8cb50bc87c0b1c
β βββ 931c77a740890c46365c7ae0c9d350ba3cca908f
β βββ e76620f83d5f5b69efd3d87e3dc180c1bd21df9fbebacfd4335e5e1efcc018da
βββ refs
β βββ main
βββ snapshots
βββ 0a363e9161cbc7ed1431c9597a8ceaf0c4f78fcf
βββ config.json -> ../../blobs/0351d1d6870005e865747b781b5d7c23ea0459cd
βββ model.bin -> ../../blobs/e76620f83d5f5b69efd3d87e3dc180c1bd21df9fbebacfd4335e5e1efcc018da
βββ preprocessor_config.json -> ../../blobs/931c77a740890c46365c7ae0c9d350ba3cca908f
βββ tokenizer.json -> ../../blobs/17456db595adc78a973f97d69d8cb50bc87c0b1c
βββ vocabulary.json -> ../../blobs/0adcd01e7c237205d593b707e66dd5d7bc785d2d
5 directories, 11 filesIt's best to specify directly:
# ~/models/<model-name>/ is the local path
--local-dir ~/models/<model-name>/The resulting directory structure is basically the same as the model repository itself, making it easier to inspect, package, upload, and later load directly.
Note that currently, even when using --local-dir, Hugging Face will create a .cache/huggingface/ metadata directory in the target directory, used to determine whether files need to be re-downloaded.
After the model download completes and you've confirmed you no longer need to sync updates, this metadata directory can be deleted.
ModelScope CLI
You need to use the ModelScope SDK to download models. Official documentation
Installation
pip install -U modelscope
# Confirm the installation succeeded
modelscope --versionUsage
ModelScope's logic is very similar to Hugging Face's. The basic command-line download format is:
modelscope download --model <model-id> --local_dir <directory>The parameter naming is slightly different between the two:
Hugging Face --local-dir
ModelScope --local_dirOne uses -, the other uses _.
- Default download location:
~/.cache/modelscope/hub. To change it, set the environment variableMODELSCOPE_CACHE. On macOS, add the following to~/.zshrc:
export MODELSCOPE_CACHE=/path/to/your/cache-
<model_id>is the model ID, obtained from the model's homepage, consisting of the organization name and model name, e.g.Qwen/Qwen-Image-2.1; -
Downloading private models requires login;
-
Download the entire repository to a directory:
modelscope download --model Qwen/Qwen-Image-2.1 --local_dir /path/to/your/models/- Download a single file to a directory:
modelscope download --model Qwen/Qwen-Image-2.1 README.md --local_dir /path/to/your/models/- To download multiple specified files, write the file names one after another:
modelscope download \
--model <model-id> \
README.md \
model_index.json \
--local_dir /path/to/your/models/- Use
--includeto select the files you need according to a glob pattern:
modelscope download \
--model <model-id> \
--include "*.safetensors" \
--local_dir /path/to/your/models/You can also specify multiple patterns:
modelscope download \
--model <model-id> \
--include "*.json" "*.safetensors" \
--local_dir ./models/model-name/- Use
--excludewhen you need most of the files but want to exclude certain ones. For example, if a repository provides both.safetensorsand.binweights and you only need safetensors, you can directly exclude.bin:
modelscope download \
--model <model-id> \
--exclude "*.bin" \
--local_dir /path/to/your/models/- Parameters:
| Parameter | Short | Type | Default | Description |
|---|---|---|---|---|
repo_id | - | str | - | Positional arg, repository ID (optional, can also be set via --model) |
files | - | str | - | Positional arg, specifies the files to download (supports multiple) |
--model | - | str | None | Model ID (mutually exclusive with --dataset) |
--dataset | - | str | None | Dataset ID (mutually exclusive with --model) |
--repo-type | - | choice | model | Repository type (model/dataset), used with the positional repo_id |
--revision | - | str | None | Version / branch / tag |
--cache_dir | - | str | None | Cache directory |
--local_dir | - | str | None | Local directory (takes precedence over cache_dir) |
--include | - | list | None | File glob patterns to include |
--exclude | - | list | None | File glob patterns to exclude |
--token | - | str | None | Access token (required for private models) |
--endpoint | - | str | None | ModelScope service endpoint |
--max-workers | - | int | default | Maximum number of concurrent download threads |
BitaHub CLI
BitaHub provides bita client versions for Windows, Linux, macOS Intel, and macOS Apple Silicon. I use macOS Apple Silicon. After installation, confirm:
bita --helpBitaHub officially recommends that on macOS you ensure /usr/local/bin/ exists before installation, then install the .pkg for your architecture.
mkdir -p /usr/local/bin/Once the model is prepared, you can upload it to BitaHub. First, you need to create a file system in BitaHub's File Storage.
For models of tens of GB, using browser upload directly is not recommended. BitaHub also officially recommends using the bita command-line fast-upload tool when a single file exceeds 10GB, when there are many files, or when you need to upload from a local terminal.
- Log in:
bita login -e https://www.bitahub.comThe CLI uses browser-based authorization. The terminal displays an authorization URL and a one-time authorization code. Complete the authorization in the browser β you don't need to enter your account password directly in the terminal.
- Get the upload command:
Go to BitaHub's Storage Management page, select the corresponding file system, and click the Upload button. The page will automatically generate something like:
bita upload \
-e https://www.bitahub.com \
-b <bucket-id> \
-o / \
-l <local file or folder path>Where:
-e BitaHub platform address
-b bucket id corresponding to the current file system
-o target directory in the file system for upload
-l local file or folder pathIn practice, you don't need to manually enter the bucket-id β it's safest to directly copy the command generated by BitaHub's current upload page.
Mount and Run the Model
First, complete all the preparation work:
Local model download complete
β
File system created
β
Model uploaded
β
BitaHub model createdAfter everything is prepared, then create the GPU development environment.
The model deployment workflow officially provided by BitaHub is also: upload models to the file system first, create models from these files, and finally mount the corresponding model when creating the development environment.
When creating the development environment, select the previously created file system or model under Storage Mount. After the model is mounted, you can read the model files directly from the mount directory inside the container, for example:
/oceanz/models/Qwen/Qwen-Image-2.1The actual path depends on the name used when creating the model and the mount configuration. Afterwards, the code can load using the local path directly, for example:
pipe = QwenImage21Pipeline.from_pretrained(
"/oceanz/models/Qwen/Qwen-Image-2.1",
torch_dtype=torch.bfloat16,
local_files_only=True
)For GGUF with llama.cpp:
llama cli -m /oceanz/models/Qwen/Qwen3.8-27B-Uncensored/Qwen3.8-27B-Uncensored-Q5_K_M.ggufIf it's a vision model, also mount the corresponding multimodal projector:
llama cli -m /oceanz/models/Qwen/Qwen3.8-27B-Uncensored/Qwen3.8-27B-Uncensored-Q5_K_M.gguf --mmproj /oceanz/models/Qwen/Qwen3.8-27B-Uncensored/mmproj-Qwen3.8-27B-Uncensored-F16.ggufThis way, once the GPU instance starts, you can basically begin environment setup and inference directly, without waiting for tens of GB of model downloads.
Related Links
-
If you're planning to register for BitaHub, you can use my invitation link: BitaHub Invitation Registration
-
Or visit directly: BitaHub