Author's Description
Shisa V2 Llama 3.3 70B is a bilingual Japanese-English chat model fine-tuned by Shisa.AI on Meta’s Llama-3.3-70B-Instruct base. It prioritizes Japanese language performance while retaining strong English capabilities. The model was optimized entirely through post-training, using a refined mix of supervised fine-tuning (SFT) and DPO datasets including regenerated ShareGPT-style data, translation tasks, roleplaying conversations, and instruction-following prompts. Unlike earlier Shisa releases, this version avoids tokenizer modifications or extended pretraining. Shisa V2 70B achieves leading Japanese task performance across a wide range of custom and public benchmarks, including JA MT Bench, ELYZA 100, and Rakuda. It supports a 128K token context length and integrates smoothly with inference frameworks like vLLM and SGLang. While it inherits safety characteristics from its base model, no additional alignment was applied. The model is intended for high-performance bilingual chat, instruction following, and translation tasks across JA/EN.
Key Specifications
Supported Parameters
This model supports the following parameters:
Performance Summary
Shisa V2 Llama 3.3 70B demonstrates a balanced performance profile, excelling in specific areas while maintaining competitive standing in others. Its speed performance is moderate, ranking in the 38th percentile across benchmarks. However, it consistently offers among the most competitive pricing, ranking in the 92nd percentile. The model exhibits exceptional reliability, boasting a 97% success rate, indicating minimal technical failures. In terms of benchmark performance, Shisa V2 shows remarkable strength in Instruction Following, achieving perfect accuracy in one instance and ranking among the top models for both accuracy and speed. It also achieves perfect accuracy in Email Classification, proving to be highly accurate and cost-effective in this domain. Its Hallucinations (Baseline) accuracy is strong at 92.0%, suggesting a good ability to acknowledge uncertainty. General Knowledge and Ethics performance are solid, ranking in the 56th and 63rd percentiles respectively. However, the model shows weaknesses in more complex tasks such as Reasoning (36th percentile) and Coding (29th percentile), and a second Instruction Following benchmark also showed a lower accuracy of 50.5%. These areas suggest opportunities for further optimization.
Model Pricing
Current Pricing
| Feature | Price (per 1M tokens) |
|---|---|
| Prompt | $0.05 |
| Completion | $0.22 |
Price History
Available Endpoints
| Provider | Endpoint Name | Context Length | Pricing (Input) | Pricing (Output) |
|---|---|---|---|---|
|
Chutes
|
Chutes | shisa-ai/shisa-v2-llama3.3-70b | 32K | $0.05 / 1M tokens | $0.22 / 1M tokens |
|
Chutes
|
Chutes | shisa-ai/shisa-v2-llama3.3-70b | 32K | $0.05 / 1M tokens | $0.22 / 1M tokens |
Benchmark Results
| Benchmark | Category | Reasoning | Strategy | Free | Executions | Accuracy | Cost | Duration |
|---|