Tool overview
Inferless is listed under AI Infrastructure & MLOps AI tools.
What is Inferless?
Inferless deploys machine-learning and generative models as serverless GPU endpoints. It builds containerized runtimes, scales replicas with demand, supports fractional and dedicated Nvidia GPUs, and charges for healthy running time by the second.
Best for
Developers deploying bursty GPU model endpoints
Who is it for?
Decision note
Rebuilt from the original export under the complete V412/V411 factual-source-verification workflow. Preview only. Apply remains blocked until Failures = 0, Warnings = 0, Unmapped = 0, Missing = 0; explicit clears are reviewed; image import is disabled or V380 accepts the asset; a representative WordPress edit screen is compared with the export and proposed row; and a post-Apply zero-change Preview succeeds.
Key features
Per-second serverless GPU billing
Fractional and dedicated T4, A10, and A100 GPUs
Scale-to-zero endpoints and configurable concurrency
Custom Docker runtimes and model-repository workflows
Use cases
Deploy generative and ML model APIs
Serve image, language, audio, and vision models
Run bursty GPU inference workloads
Replace continuously provisioned GPU endpoints
Pros
- Starts at $0.33 per GPU hour
- Promotional free compute credit without a card
- No charge when minimum replicas are zero and no worker runs
Limitations
The pricing page presents both a ten-hour statement and a $30 promotional-credit statement
the current dashboard grant should be confirmed at signup. Security isolation does not remove the need to protect model code, data, and API credentials.
Pricing details
Billing options
Pricing note
Shared T4 starts at $0.000092/second or $0.33/hour. Shared A10 is $0.61/hour and shared A100 is $2.68/hour. The page promotes free starter compute and $30 credit; the exact active signup grant should be confirmed in the dashboard.
Supported languages
- English
Integrations
Hugging Face models
GitHub repositories
Docker containers
AWS storage workflows
Custom runtimes
Please log in to join the discussion.