Skip to content
KaminClarity for your next step.
← Back to Kamin

What is decided, and what remains research?

This page separates prototype decisions from research results. No validation study outcomes or new accuracy figures are presented here; this is a reviewable protocol outline only.

01
Resolved for prototype

First reference department

The Department of Information Technology, Faculty of Computing & Information Technology, King Abdulaziz University is the current course namespace and pilot-adapter reference. This does not imply institutional approval or official launch.

02
Resolved for public release

No generative Arabic model drives current Fit judgments

The public release uses in-browser Arabic OCR (Tesseract language data) for text extraction, followed by governed rules, ontology mappings, and graph traversal. Any future Arabic language model is a separate decision requiring evaluation, licensing, and governance before affecting product judgments.

03
Plan — no results

Retrospective study

Goal: compare Kamin outputs on a de-identified historical record set against a pre-defined human reference. Sample, inclusion rules, reference definition, outcomes, and analysis plan must be frozen before data access.

04
Plan — no results

Evidence-deposit completion rate

Goal: measure the share of invited participants who complete an approvable evidence deposit and reach a first capability profile. This metric must remain separate from optional research consent and must not collect transcript content in telemetry.

Measurement principles

  • Published methodology Targets are not Results.
  • Accuracy evaluation requires a documented human reference.
  • Mapping coverage, evidence-cited judgment ratio, and time-to-first-value are product-health metrics.
  • Psychometrics and sensitive traits do not enter Fit before independent validation.

Ethics status

This page does not claim ethics/IRB approval for these studies. Any use of real records or publication of results requires the appropriate ethical and regulatory pathway before execution.

Proposal dated 30 September 2026 — no results

Nine-month validation roadmap

Month 1 starts after protocol and institutional approval, not when this page is published. Dates and responsibilities require supervisor confirmation; no partnerships or recruited samples are announced.

The semester goal is a usable product and an initial study in weeks 10–12 after approval. Later months deepen validation and follow-up, with final handover in month ten. This roadmap does not postpone the semester deliverable.

Month 1–2: Scope and protocol approval

Supervisor + participating institutions · Confirm colleges, approvals, measures and data responsibilities. Target: two specialists review at least 20 pathways before mappings are accepted.

Month 2–3: Formative usability

UX team + advisors · Target: 12 Saudi students over two rounds, including Arabic and mobile. Measure six tasks using synthetic records and resolve blockers before scaling.

Month 3–4: Bounded end-of-semester pilot

Supervisor + approved colleges · Target: 60 participants, 20 per college. Report independent completion, assistance and failure with uncertainty, separated by language and device.

Month 5–6: Issuer-verification prototype

Evidence team + authorized issuer · Target: 20 synthetic test cases covering invalid signatures, expiry and revocation. No institutional evidence upgrades before negative cases and security review pass.

Month 7–8: Follow-up and reevaluation

Team + advisors · Approved optional follow-up, comparing improvements with the baseline on a separate test set; no causal employment-effect claims.

Month 9: Continuation and pricing decision

Supervisor + career center · Target: 5 interviews with purchasing decision-makers; review results, limitations and operating cost. No automatic transition to a paid service.

Proposed pre-pilot thresholds — not results

  • Independent completion ≥80% per core task; failures and assisted attempts remain in the denominator, and withdrawals are reported.
  • Source-and-limit comprehension (2 of 2) in ≥80% of eligible attempts; two reviewers code a prespecified subset.
  • Zero false evidence upgrades in the negative verification suite and zero unresolved critical privacy defects before a real-record pilot.
  • Freeze definitions, sample and stopping rules before the study. These are continuation targets, not proof of accuracy, fairness or improved employment.

Usability tasks: build a profile, correct evidence, understand a recommendation, choose a step, restore a backup, and correct, withdraw and delete data. Report counts and uncertainty alongside rates; a missed target requires remediation and remeasurement.

Review speed and quality — pre-study targets

Uncalibrated proposal; no published field results. Freeze definitions, cases and targets with the supervisor before collecting results; separate language, device and reviewer role.

Time to confident acceptance or rejection

From first presentation of the review task and recommendation to acceptance or rejection with the highest of three confidence levels and a correct source-and-limit explanation assessed independently. Initial target: ≥80% of eligible tasks end in a qualifying decision, with median time ≤120 seconds among qualifying attempts. Selecting “confident” alone does not make a decision correct.

Report contests, unresolved outcomes, failures, assistance, interruptions and withdrawals separately; do not remove them to improve completion rates. The optional local timer starts when explicitly enabled, so it is not automatically equivalent to study first-exposure timing. Interrupted local attempts are excluded only from time summaries, with counts disclosed.

Errors that escaped verification

In a frozen synthetic suite: known-defective cases accepted after review divided by all known-defective cases presented for review. Initial target ≤5%, with zero accepted false evidence upgrades or critical privacy defects. Also report confirmed errors in an independent sample of accepted decisions with its denominator; do not combine that rate with the seeded-suite rate. Untriaged contests are not confirmed errors.

Rule and review-boundary regression checks now run in the repository. Student/advisor decision studies, institutional-source verification and central contest triage have not been completed. Engineering test results are not human-study results.

Task-loop measures

Pre-study targets, not results: share of started tasks that reach an attached output; share of outputs meeting the rubric under an independent reviewer; gap between self-assessment and reviewer assessment; and the number of tasks assessed by the two HR and domain experts before adoption (target: all six).

Research appendix: traceability identifiers

D-05: reference scope · D-01: model choice · H3: retrospective validation · H4: evidence-deposit completion.