
Isaac Emevo
Verified Expert in Support and Operations
Support & Operations Expert
Smyrna, GA, United States
Toptal member since March 9, 2026
Isaac is a data annotation specialist and GenAI expert with 5+ years of experience in AI data services, end-to-end refinement of large language models (LLMs). He evaluates multimodal models by ranking hundreds of text, audio, video, and image outputs daily to ensure alignment and coherence. Isaac creates high-quality datasets and rigorous rubrics for systematic comparative judgment, deep QA reviews, ML data audits, and precise annotation guidelines for complex projects.
Expertise
- Data Annotation
- Generative Artificial Intelligence (GenAI)
- Language Models
- Natural Language Processing (NLP)
- Prompt Engineering
- Quality Assurance (QA)
- Rubrics
- Spreadsheets
Work Experience
Data Partner | Generalist
TELUS Digital
- Tagged, categorized, and reviewed text, images, and documents.
- Gathered and organized high-quality datasets for AI training.
- Reviewed and validated data for precision and consistency.
- Worked with an international team to hit project milestones.
Generalist
Mercor
- Running prompts through models and assessing the quality, accuracy, and coherence of the results.
- Comparing AI-generated responses to determine which is superior, which aids in training models to reason more effectively.
- Detecting subtle errors, inconsistencies, or biases in AI responses while adhering to or refining detailed evaluation guidelines.
- Developing questions and answers and, in some cases, video, text, image, and audio-based evaluations for AI training, especially for specialized, short-term, or weekend-only projects.
AI-Trainer
- Evaluated prompts and tasks for review on general tasks within Software Engineering.
- AI Trainer Projects. Did not do coding- or software-specific tasks, just generalist reviews and evaluations.
- Led to accuracy in how ai models review documents.
STEM Domain Task Creation (Advanced Level)
Pasiflora AI
- Delivered original STEM prompts at the graduate to PhD level.
- Delivered anonymized financial documents and data exports.
- Led to the improvement in Model performance for an AI Lab.
Audio Model Trainer | Digital Annotation Expert
Mercor
- Performed daily preference ranking of hundreds of multimodal AI outputs, including images, audios, and videos, against prompts to evaluate model performance and alignment.
- Applied systematic comparative judgment techniques to assess relevance, quality, and coherence of generative model outputs, ensuring consistent evaluation across diverse modalities.
- Maintained high accuracy and throughput in large-scale annotation workflows, contributing to reliable training data for model fine-tuning and preference optimization.
- Collaborated with cross-functional AI teams to identify quality trends and provide actionable feedback for improving multimodal generative capabilities.
- Utilized specialized labeling platforms and guidelines to ensure precise annotation, supporting preference-based reinforcement learning and human feedback loops.
Senior Data Quality Specialist, Safety
Cohere
- Labeled, ranked, audited, and corrected annotator and ML LLM data.
- Recommended optimization opportunities through human-in-the-loop development and operations.
- Completed text-based tasks efficiently and attentively, and interacted with Tier 1 and Tier 2 data.
- Provided feedback to cross-functional team members.
Data Annotator | Writer
Innodata
- Developed and implemented retrieval augmented generation (RAG) models, enhancing the software automation capabilities of LLMs.
- Reviewed and analyzed paragraphs of text or responses generated by AI models, such as chatbots.
- Labeled both the prompts and the AI-generated responses according to specific guidelines.
- Evaluated and rated the AI responses based on various criteria, including helpfulness, accuracy, and grammatical correctness.
- Handled sensitive or potentially offensive content responsibly, including but not limited to light-toxic materials involving violence, identity-based attacks, and sexual harassment.
- Facilitated the training of generative AI models on various client projects, meticulously crafting prompts and responses in adherence with industry-recognized best practices for high-quality prompts.
- Utilized AWS and in-house proprietary tools effectively, contributing to the seamless execution of project tasks.
Shopper Analyst | eCommerce Analyst Associate Tier 1
TELUS International
- Produced insightful statistics on online sales proactively.
- Identified and analyzed patterns in consumer purchases.
- Assessed shifts in the online retail industry to provide valuable information to advertising managers.
- Collaborated with developers to tailor online transaction procedures, resulting in a seamless online shopping experience for customers.
- Utilized various Google Suite tools and Workday to excel in the role.
Document Processor | Register
Stride Learning
- Verified the accuracy of all information provided before entering it into the system.
- Maintained detailed records of client interactions and outcomes.
- Entered data efficiently and accurately into the designated database.
- Organized files in a systematic manner for easy access and retrieval.
- Filed documents in accordance with established protocols and standards.
Tasker Success Manager
Scale AI
- Served as the primary point of contact for new taskers.
- Provided onboarding and training to new taskers, including, but not limited to, explaining the platform, best practices for task completion, project-specific guidance, and clarifying any doubts.
- Monitored tasker performance and provided feedback as needed.
- Answered questions and provided support via email, chat, or video calls.
- Maintained a positive and friendly attitude in all interactions with taskers.
Project History
Dataset Creation and Evaluation (Finance)
Created and reviewed financial dataset for LLM at STEM and generalist levels.
I developed a financial dataset and benchmarking framework to evaluate LLMs across STEM quantitative computation and generalist domain reasoning. For the STEM tier, I curated multi-step mathematical workflows—including derivative pricing, portfolio optimization, time-series forecasting, and code-based modeling—to assess numerical precision and algorithmic soundness. For the generalist tier, I designed tasks covering 10-K/10-Q synthesis, earnings call sentiment analysis, regulatory compliance, and advisory dialogues to evaluate macro-level financial literacy and strategic context.
I implemented a dual-pipeline evaluation methodology to ensure data fidelity. Quantitative outputs were validated via deterministic programmatic verification against exact ground-truth solutions under strict error tolerances. Qualitative, reasoning-heavy responses were assessed using structured rubrics combining automated LLM-as-a-judge scoring with expert human review for factual consistency. By bridging computational finance with commercial reasoning, this benchmark reduced hallucinations, measured domain proficiency, and accelerated reliable financial AI adoption.
Project Marvel
Constructed and labeled extremely complex visual question answering datasets from academic images to develop state-of-the-art AI models that can perform multi-step reasoning.
The goal of the project was to develop domain-specific datasets to train advanced Large Language Models (LLMs) to effectively interact with visual content from specific domains. I was a creator and reviewer on this initiative to analyze complex textbook images, charts, and diagrams to formulate non-extractive prompts requiring high-level cognitive effort and visual interpretation. I created detailed explanations with a strict “Observe, Explain, Answer” framework to trace out logical, line-by-line deductions that demonstrate expert domain knowledge. Additionally, I conducted extensive quality validation audits on visual question answering (VQA) tasks, evaluating for scientific accuracy, coherence, complexity, and strict format compliance to maximize data integrity for AI training.
Education
Associate's Degree in Business Administration
University of the Commonwealth Caribbean - Kingston, Jamaica
Certifications
Claude with the Anthropic API
Anthropic
Claude Code in Action
Anthropic
LLM Graded Certification
Appen
AI/Machine Learning Expert
RWS Group
Skills
Tools
Slack, Asana, Jira, FullStory
Administrative Operations
Quality Assurance (QA), Process Documentation, Data Entry, Document Control, File Management, Bookkeeping
Customer Support
Communication
Productivity Suites
Google Sheets, Google Docs, Spreadsheets
Professional Skills
Data Annotation
Quality & Performance
Implementing QA Feedback
Sales Operations
CRM
Collaboration Tools
Zoom
CRM & Sales Tools
Salesforce
Operations
Databases
Skill Ladders
Performance Tracking, Onboarding
Commerce Platforms
Shopify
HR & Recruiting Tools
Workday
Product Support
API Tokens
Project Management Tools
Notion
Other
Hubstaff, Google Workspace, Amazon SageMaker, Lablebox, Computer Vision, Generative Artificial Intelligence (GenAI), Artificial Intelligence (AI), Agentic AI, Rubrics, Data Modeling, Data Collection, A/B Testing, Reinforcement Learning from Human Feedback (RLHF), Prompt Engineering, Large Language Model Operations (LLMOps), Creative Writing, Machine Learning, Red Teaming, Language Models, AI Prompting, Data Analysis, Data Visualization, SFT, Digital Media, Multimedia, Multimodal, Model Stumping, Natural Language Processing (NLP), Technical Support, Data Analytics, Customer Support, QA Testing, Retrieval-augmented Generation (RAG), Linguistics, Human-in-the-Loop Machine Learning, Learning Management Systems (LMS), Project Management, Business Intelligence (BI), Business Strategy, Administrative Assistance, Business Law, Economics, Technical Writing, Data Research, Technical Training, Data Management, Search Engine Optimization (SEO), SQL, Microsoft SQL Server, Microsoft Office, Amazon Web Services (AWS), Data Science, Accounting, Politics, JSON, AI Research, Okta, Google Ads, SDKs, eCommerce, Database Management, Total View Enrollment, Workday Payroll, Python, APIs, Model Context Protocol, GitHub, Claude Code, Terminal Servers, Record Keeping, Financial Data, Finance Strategy, Financial Planning & Analysis (FP&A), Mechanical Engineering, Large Language Models (LLMs), STEM, Biology, Physics, Engineering, QA Documentation
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring