C

Data Scientist

Coforge · Pune District, Maharashtra

3–10 yrs experiencefull_timePosted 1w ago

Job description

Job Title: Data Scientist Skills: Artificial intelligence, Machine Learning, NLP, Gen AI, Python, Rest API, Agentic AI, LLM, RAG, Devops and AWS Experience: 4+ years Location: Pune and Hyderabad Duration: Full time We at Coforge are hiring for Data Scientist role with following skill sets: **LLM & Generative AI** - Design, build, and deploy **LLM-powered applications** using frameworks such as **LangChain** , **LlamaIndex** , or OpenAI API. - Develop and optimize **prompt engineering** strategies (few-shot, chain-of-thought, RAG) to improve the accuracy, consistency, and reliability of LLM outputs. - Implement **Retrieval-Augmented Generation (RAG)** pipelines using vector databases (e.g., FAISS, Pinecone, Chroma, Weaviate). - Fine-tune pre-trained LLMs (e.g., GPT, LLaMA, Mistral, Falcon, Claude,Gemini) on domain-specific datasets. - Validate and structure LLM outputs using **Pydantic** models and output parsers to ensure data integrity. **Natural Language Processing (NLP)** - Build end-to-end NLP pipelines for real-world tasks including: - **Named Entity Recognition (NER)** - **Text Classification & Sentiment Analysis** - **Information & Data Extraction from Documents** - **Document Summarization & Question Answering** - **Semantic Search & Document Similarity** - Work with the **Hugging Face Transformers** ecosystem to leverage and fine-tune pre-trained models (BERT, RoBERTa, T5, etc.). - Process large-scale unstructured text data from various sources such as PDFs, emails, scanned documents (OCR), and web content. **Anomaly Detection** - Design and implement anomaly detection systems for various domains, including: - **Financial fraud detection** (unusual transactions, payment anomalies). - **Operational anomalies** (system logs, network traffic, sensor data). - **Text-based anomalies** (unusual document patterns, suspicious NLP signals). - Apply a wide range of anomaly detection techniques including: - **Statistical Methods:** Z-score, IQR, CUSUM. - **ML-based Methods:** Isolation Forest, One-Class SVM, Local Outlier Factor (LOF). - **Deep Learning Methods:** Autoencoders, LSTM-based sequence anomaly detection, Variational Autoencoders (VAEs). - **Time-Series Methods:** ARIMA, Prophet, Seasonal Decomposition. - Build real-time and batch anomaly detection pipelines that can scale to large datasets. - Define and tune detection thresholds and alert mechanisms in collaboration with business and operations teams. **Machine Learning (ML)** - Design, train, evaluate, and deploy supervised and unsupervised machine learning models. - Perform **feature engineering** , **model selection** , **hyperparameter tuning** , and **cross-validation** . - Build and maintain end-to-end **ML pipelines** from data ingestion to model serving. - Monitor model performance in production and implement retraining strategies to address **data drift** and **model decay** . - Communicate model results, performance metrics, and business impact to technical and non-technical stakeholders. **Python & Software Engineering** - Write clean, modular, production-quality, and well-documented Python code. - Build and expose ML models as **REST APIs** using **FastAPI** or **Flask** . - Collaborate with MLOps/DevOps engineers to containerize (Docker) and deploy models in cloud environments. - Follow best practices in version control ( **Git** ), testing, and CI/CD pipelines. **Data & Analytics** - Perform **Exploratory Data Analysis (EDA)** on structured and unstructured datasets to identify patterns, trends, and anomalies. - Work with data from relational databases (SQL), data lakes, and cloud storage solutions. - Create compelling and clear **data visualizations** (Matplotlib, Seaborn, Plotly) to communicate findings.