Sign language translator app
Publicada el 2026-07-20
Descripción de la oferta
We are building a bidirectional, real-time web application that translates between spoken English and Nigerian Sign Language (NSL) using an animated 3D avatar. The platform will support multiple use cases including live conversations, website/video accessibility, and extended hardware integration. Core Features 1. Bidirectional Real-Time Translation • Signer Mode: User signs in NSL to the camera → AI recognizes signs → outputs natural English text + speech. • Speaker Mode: User speaks English into the microphone → AI translates to NSL → 3D animated character signs in real time. 2. Additional Key Features • Captions Tool: Real-time captions for spoken content and sign language interpretation. • Live Translation on Websites & Embedded Videos: Embeddable widget for live translation on any website or video player (captions + avatar signing). • Embedded API: Public API for developers to integrate the translation service into third-party apps and platforms. • Smart Glasses Support: Optimized output mode for smart glasses (low-latency pose data stream or simplified avatar view). • Custom Avatar Upload: Users can upload their own avatar/character, which the system will rig and use as the signing avatar. 3. Animation & Output Requirements • Real-time generative skeleton-driven animation (Transformer/LSTM-based or equivalent) targeting sub-second latency and 30 FPS fluid motion. • Support for custom uploaded avatars: automatic rigging, bone mapping, and pose application (Quaternions/Euler angles). • Smoothing, co-articulation, facial expressions, and non-manual markers essential for natural NSL. 4. Data & Model Strategy • No pre-trained NSL model exists yet. • Plan to collect ~1,000 videos of diverse NSL signers for training both recognition and generation models. • Design the system with transfer learning/fine-tuning in mind for the custom dataset. 5. Technical Priorities • Fully web-based, low end-to-end latency (<1–2 seconds). • MediaPipe (or equivalent) for real-time pose estimation. • Privacy-focused, on-device inference where possible. • Responsive UI with camera/mic access and clean embedding options. • Extensible architecture for future languages and features.
Skills
Fuente original: freelancer