Description review

Gmail & Google Calendar AI Assistant Evaluator

micro1 · United States · back to the listing

HR standards

63/100

needs work

Title ↔ description

62/100

needs work

Reads as

Unclear

no confident match

What this role officially is

software tester — ESCO, the EU occupation classification

Software testers perform software tests. They may also plan and design them. They may also debug and repair software although this mainly corresponds to designers and developers. They ensure that applications function properly before delivering them to internal and external clients.

Also known as: application software tester, unit tester, application tester, software application tester, tester, module tester

How others title the same work

Large employers

  • Senior Software Quality Engineer Adobe
  • Distributed Systems Testing Software Engineer, Python / Go Canonical
  • Ubuntu Linux Kernel Test Engineer Canonical
  • Distributed Systems Testing Software Engineer, Python / Go Canonical Ltd.
  • Ubuntu Linux Kernel Test Engineer Canonical Ltd.

Startups

  • Software Engineer, QA & Test Automation AviaryAI

What the listing never says

  • 18 bullet points. Long requirement lists deter qualified candidates, who read them as hard gates. Scope clarity

The listing, marked up

Nothing in the wording of this listing tripped a check. The scores above still judge how complete and coherent it is.

Role Title: Gmail & Google Calendar AI Assistant Evaluator

Role Type: Contractor, Remote

Location: United States only

micro1 is engaging Gmail & Google Calendar AI Assistant Evaluators to collaborate on a customer-driven project enhancing AI assistant quality through real-world task evaluation. In this role, you'll use your everyday experience managing your digital life — email, calendars, bookings, and coordination — to help train and evaluate next-generation AI assistants. Your work will shape how AI models handle the real tasks people delegate every day.

Key Responsibilities:

• Complete realistic, everyday digital tasks (emailing, scheduling, booking, coordinating) using AI assistant tools under standardized, repeatable conditions.

• Give an AI assistant clear instructions for real-world tasks, then carefully review and verify its work for accuracy and completeness.

• Apply detailed grading guidelines to evaluate, rank, and annotate assistant responses, providing structured feedback that improves model performance.

• Test scenarios involving email threads, calendar invites, online bookings, and multi-person coordination, ensuring coverage of diverse situations and edge cases.

• Utilize AI training tools and platforms to record findings and submit evaluation data in alignment with project standards and milestones.

• Collaborate with project trainers and contributors to resolve ambiguities, refine guidelines, and ensure evaluation consistency across the team.

• Participate in ongoing quality reviews, incorporating feedback to maintain rigorous evaluation standards.

Required Skills and Qualifications:

• Are based in the US, have an iPhone with iMessage, and have a Facebook or Instagram account.

• Daily, hands-on use of Gmail and Google Calendar to run your own life: sending emails, accepting invites, scheduling, and rearranging plans.

• Regular experience coordinating with others over email or text — setting up meetups, following up on threads, or confirming shared plans and payments.

• Comfort booking things online independently, such as restaurants, appointments, travel, or deliveries.

• Experience using AI tools for real, practical tasks, with comfort giving an AI assistant instructions and checking its work.

• Exceptional attention to detail and accuracy when reviewing AI-generated output.

• Strong written communication skills for clear reporting and timely updates on progress and challenges.

• Demonstrated time management and self-organization for independent, remote project participation.

Preferred Qualifications:

• Prior experience with AI evaluation, data labeling, user testing, or similar feedback-driven projects.

• Familiarity with prompt writing and a practical understanding of AI assistant behaviors and limitations.

• Experience grading or reviewing AI outputs for accuracy, helpfulness, and safety.

Originally posted on Himalayas

How this was produced

Highlights are found by rule, not by a model: each one is a phrase matched at a known position, and every note is a template we wrote. The two scores come from a typed-decision model (Jev) that reads the listing against the official role definition and real listings for the same role, and returns probabilities rather than prose — it never writes any of the words on this page, and never chooses what to highlight.

Deterministic penalty applied to the HR score: 4 points (from 67 before penalties). Reviewed 26 Sep 2026.