Description review
Professionals: Designing Challenging AI Prompts
Terac · United States · back to the listing
HR standards
68/100
needs work
Title ↔ description
40/100
poor
Reads as
Unclear
no confident match
What this role officially is
software tester — ESCO, the EU occupation classification
Software testers perform software tests. They may also plan and design them. They may also debug and repair software although this mainly corresponds to designers and developers. They ensure that applications function properly before delivering them to internal and external clients.
Also known as: application software tester, unit tester, application tester, software application tester, tester, module tester
How others title the same work
Large employers
- Senior Software Quality Engineer Adobe
- Distributed Systems Testing Software Engineer, Python / Go Canonical
- Ubuntu Linux Kernel Test Engineer Canonical
- Distributed Systems Testing Software Engineer, Python / Go Canonical Ltd.
- Ubuntu Linux Kernel Test Engineer Canonical Ltd.
Startups
- Software Engineer, QA & Test Automation AviaryAI
What the listing never says
- No pay range published. Candidates cannot tell whether applying is worth their time. Pay transparency
- No location or timezone policy stated, so a candidate cannot tell where they may work from. Scope clarity
The listing, marked up
Nothing in the wording of this listing tripped a check. The scores above still judge how complete and coherent it is.
What We're Researching
We're running a paid study to build a bench of people who are exceptionally good at designing tasks that expose AI model limitations. By turning real-world workflows into demanding requests, we can better evaluate where current models break down. This initial trial helps us identify individuals suited for ongoing prompt engineering and evaluation work.
How It Works
You will spend about an hour translating a complex workflow from your job or personal life into a demanding prompt that requires reasoning and real-world lookup. After running it in ChatGPT to identify where the model fails, you will refine the prompt until it breaks the system. Finally, you will write a clear grading rubric that a stranger could use to evaluate any AI's attempt at your task. This entire process is screen-recorded, as we are assessing your thought process just as much as the final submitted files.
Who This Is For
We welcome professionals, domain experts, and power users who have deep knowledge of specific workflows. You need to be capable of evaluating an AI's output within seconds and comfortable working on a laptop or desktop with a ChatGPT account. Candidates who excel at this trial will be considered for a long-term bench of evaluators.
What You'll Do
• Pick a familiar workflow and convert it into a demanding AI prompt
• Test your prompt in ChatGPT to find failure points, making it harder if the AI succeeds
• Write a comprehensive rubric for grading the AI's performance
• Share your screen, camera, and microphone while completing the task
• Submit your prompt, failure notes, rubric, and the generated output file
Who Should Apply
• Deep familiarity with a specific professional or personal workflow
• Ability to quickly evaluate the accuracy and quality of AI outputs
• Access to a laptop or desktop computer
• An active ChatGPT account
• Comfortable being screen-recorded while thinking through complex tasks
Compensation
$20 one-time
Ready to participate?
Start your paid interview now
About Terac
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Learn more at or on YouTube at @jointerac.
Originally posted on Himalayas
We're running a paid study to build a bench of people who are exceptionally good at designing tasks that expose AI model limitations. By turning real-world workflows into demanding requests, we can better evaluate where current models break down. This initial trial helps us identify individuals suited for ongoing prompt engineering and evaluation work.
How It Works
You will spend about an hour translating a complex workflow from your job or personal life into a demanding prompt that requires reasoning and real-world lookup. After running it in ChatGPT to identify where the model fails, you will refine the prompt until it breaks the system. Finally, you will write a clear grading rubric that a stranger could use to evaluate any AI's attempt at your task. This entire process is screen-recorded, as we are assessing your thought process just as much as the final submitted files.
Who This Is For
We welcome professionals, domain experts, and power users who have deep knowledge of specific workflows. You need to be capable of evaluating an AI's output within seconds and comfortable working on a laptop or desktop with a ChatGPT account. Candidates who excel at this trial will be considered for a long-term bench of evaluators.
What You'll Do
• Pick a familiar workflow and convert it into a demanding AI prompt
• Test your prompt in ChatGPT to find failure points, making it harder if the AI succeeds
• Write a comprehensive rubric for grading the AI's performance
• Share your screen, camera, and microphone while completing the task
• Submit your prompt, failure notes, rubric, and the generated output file
Who Should Apply
• Deep familiarity with a specific professional or personal workflow
• Ability to quickly evaluate the accuracy and quality of AI outputs
• Access to a laptop or desktop computer
• An active ChatGPT account
• Comfortable being screen-recorded while thinking through complex tasks
Compensation
$20 one-time
Ready to participate?
Start your paid interview now
About Terac
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Learn more at or on YouTube at @jointerac.
Originally posted on Himalayas