We are sharing a specialised full-time consulting opportunity for US-based senior retail professionals with deep experience in merchandising, category management, buying, planning, retail operations, and structured evaluation of AI-generated outputs. This role supports a high-impact generative AI initiative focused on improving how advanced models reason through real-world retail challenges. Selected professionals will design rigorous retail tasks, evaluate model outputs against structured criteria, develop scoring frameworks, and provide practical feedback grounded in senior-level merchandising, category, planning, and operational experience. Key Responsibilities Retail Strategy & Domain Guidance Guide research and engineering teams on merchandising, category management, assortment planning, buying, and retail operations Identify gaps in model understanding across pricing, inventory, promotions, customer demand, store performance, and commercial planning Apply practical retail judgment to complex scenarios involving products, channels, suppliers, customers, and financial objectives Ensure retail tasks reflect realistic business decisions and current industry practices Retail Task & Solution Development Design challenging, domain-relevant tasks grounded in real retail practice Write accurate, well-reasoned solutions covering merchandising, category, buying, planning, and operational scenarios Develop problems requiring commercial judgment, prioritisation, data interpretation, and evaluation of competing trade-offs Ensure tasks are clear, internally consistent, and suitable for structured assessment AI Output Evaluation Evaluate AI-generated retail responses against established rubrics and scoring criteria Assess correctness, commercial judgment, reasoning quality, relevance, and practical applicability Compare alternative outputs and determine which response provides the stronger retail recommendation Identify factual errors, unsupported assumptions, weak commercial reasoning, and incomplete conclusions Provide clear written feedback that supports improvements in model behaviour Rubric Development & Quality Calibration Develop and refine evaluation guidelines for merchandising, category management, planning, and retail operations tasks Define scoring criteria covering technical accuracy, commercial judgment, reasoning, and execution feasibility Participate in calibration activities with other retail specialists Collaborate across teams to maintain consistency and accuracy throughout the training data Ideal Profile Strong candidates may have: At least 8 years of dedicated professional experience in retail Senior-level expertise in merchandising, category management, buying, planning, retail operations, or a related area Experience working within a recognised retailer, consumer brand, e-commerce platform, or comparable organisation Prior hands-on experience evaluating LLM or AI-generated outputs using structured rubrics or scoring criteria Demonstrable career progression into senior manager, director, vice president, or comparable leadership responsibilities Strong commercial judgment and the ability to assess complex retail decisions Excellent written and verbal communication skills Reliable availability for at least 35 hours per week during weekdays Educational Background A degree in business administration, retail management, merchandising, supply chain, marketing, economics, or a related field is highly relevant Graduate-level education in business, strategy, operations, analytics, or management may be helpful Equivalent senior professional experience in retail strategy or operations may also be considered Professional training in category management, merchandising analytics, procurement, or commercial planning may be valuable Nice to Have Experience working with large-scale retailers, global consumer brands, or high-growth e-commerce businesses Background in omnichannel retail, marketplace operations, inventory planning, or store operations Experience creating structured scoring rubrics, evaluation guidelines, or quality frameworks Familiarity with model training, human-feedback workflows, annotation, or AI quality assurance Strong understanding of pricing, assortment, promotions, forecasting, inventory, and margin management Experience reviewing category plans, buying strategies, merchandising proposals, or operational performance reports Previous collaboration with product, analytics, supply chain, marketing, finance, or technical teams Why This Opportunity Apply senior retail expertise to an advanced generative AI initiative Influence how AI systems reason about merchandising, category management, and retail operations Design and assess challenging retail problems grounded in practical commercial experience Collaborate with technical teams and experienced retail specialists Join a full-time remote engagement with competitive hourly compensation Contract Details Full-time W-2 contingent employment arrangement Fully remote role available to candidates based in the United States Expected commitment of at least 35 hours per week during weekdays Competitive rates between $55–$75 per hour depending on expertise and project scope Prior experience evaluating AI or LLM outputs against structured rubrics is required Applicants should clearly describe relevant AI evaluation experience in their application Immediate availability is preferred Work may include onboarding, specialty calibration, task development, and ongoing quality review Project scope and duration may be adjusted according to programme requirements and performance About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy .
Remote | AI Safety Evaluation Specialist — $55–$65/hour
24 Mag
Remote | Swedish Music & Audio Evaluation Specialist — $30–$55/hour
24 Mag
Remote Booking Assistant
Journey with Haylee
Remote Scheduling Assistant (No Experience Needed)
Journey with Haylee
RN Supervisor UM Prior Auth
Affiliates Commonspirit
TikTok Shop Growth Lead/CSM
AMZ Advisers