Overview

Junior QA Tester, LLM Evaluation & Training Jobs in New York, NY at K-Hawk Capital

Title: Junior QA Tester, LLM Evaluation & Training

Company: K-Hawk Capital

Location: New York, NY

Company Description

K-Hawk Capital is a systematic trading fund founded by Marc Preston and a diverse team of quantitative specialists with over 25 years of experience in global markets. The firm designs and trades proprietary algorithmic strategies rooted in scientific research, statistical rigor, and advanced technology. Its in-house models are supported by high-performance computing, robust infrastructure, sophisticated risk analytics, and artificial intelligence, deployed across multiple asset classes. K-Hawk Capital emphasizes long-term talent development through an apprenticeship model, offering hands-on exposure to real-time trading systems under the guidance of experienced practitioners. Team members join a collaborative environment that encourages intellectual curiosity and innovation in modern systematic investing.

K-Hawk Capital is hiring a Junior QA Tester to evaluate, stress-test, and help train our large language model. This is precision work: finding errors, inconsistencies, and failure modes that are easy to miss, documenting them so they’re actually fixable, and holding a consistent bar across a high volume of model outputs.

What you’ll do:

  • Test LLM outputs against defined prompts and scenarios, catching factual errors, logical inconsistencies, bias, or unsafe responses — including subtle failures
  • Write and execute structured test cases covering expected use cases, edge cases, and adversarial/stress scenarios
  • Log, categorize, and document issues so they’re reproducible and actionable — vague bug reports will be sent back
  • Rate and annotate model outputs consistently against internal rubrics, without drifting from the standard over time
  • Give structured, specific feedback to engineering/ML — “this feels off” is not acceptable feedback
  • Identify recurring patterns and systemic issues across large volumes of output, not just one-off errors
  • Meet daily/weekly throughput targets without sacrificing accuracy

Minimum qualifications (applications missing these will be rejected):

  • 2–4 years of demonstrated experience in QA, software testing, editorial/content review, or data annotation — recent grads must show specific, verifiable projects
  • Native or professional fluency in written English, evidenced by the application itself
  • Proven ability to work independently and hit deadlines in a remote, asynchronous environment
  • Demonstrated logical rigor — you can articulate why something is wrong, not just flag that it “seems off”
  • Direct hands-on experience with LLM tools (ChatGPT, Claude, Gemini, etc.) beyond casual/consumer use — e.g., prompt testing, red-teaming, or evaluation work
  • Experience with bug tracking or QA tooling (Jira, Linear, TestRail, or similar)
  • Background in linguistics, computer science, data annotation, or a field requiring rigorous written analysis

To apply: Provide a link to an LLM evaluation, test harness, dataset, or technical project you personally built. In 150 words or less, describe one failure the evaluation uncovered, the evidence that proved it, and the change you made, and email it all to [email protected]. Generic or hypothetical answers will not be considered.

Upload your CV/resume or any other relevant file. Max. file size: 800 MB.