{"repo":"yuniko-software/tokenizer-to-onnx-model","free":true,"listed":false,"github":"https://github.com/yuniko-software/tokenizer-to-onnx-model","clone":"git clone https://github.com/yuniko-software/tokenizer-to-onnx-model.git","description":"Convert Hugging Face tokenizers to ONNX models for cross-language compatibility (.NET, Java, Python) with embedding models","language":"Jupyter Notebook","stars":38,"topics":["csharp","dotnet","embedding-models","huggingface","inference","java","machine-learning","onnx","python","tokenization"],"license":"Apache-2.0","category":"machine-learning","readme_excerpt":"Hugging Face Tokenizer to ONNX Model ⚠️ Looking for Full BGE-M3 Functionality? This repository demonstrates basic tokenizer conversion and generates only dense embeddings . If you need the complete BGE-M3 experience with all three embedding types (dense, sparse, and ColBERT vectors), check out our new repository: BGE-M3 ONNX The new repository provides: - All three BGE-M3 embedding types (dense, sparse, ColBERT) - Production-ready implementations in C#, Java, and Python This repository demonstrates how to convert Hugging Face tokenizers to ONNX format and use them along with embedding models in multiple programming languages. Key Features - Generate embeddings directly in C#, Java, or Python without third-party APIs or services - Reduced latency with local embedding generation - Full control over the embedding pipeline with no external dependencies - Works offline without internet connectivity requirements - Cross-platform compatibility The Problem While we can easily download ONNX models from Hugging Face or convert existing PyTorch models to ONNX format for portability, tokenizers present a significant challenge. Tokenizers for embedding models are not often implemented in languages other than Python. This becomes a major obstacle when trying to use embedding models in languages like C# or Java. Developers face the difficult choice of either implementing complex tokenizers from scratch or relying on Python interop, which adds complexity and dependencies. The Solution This r","default_branch":null,"files":null,"tree":[],"storefront":"/r/yuniko-software","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/yuniko-software/tokenizer-to-onnx-model/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}