Skip to content

Cortex AI

Agentic AI assistant for PCI DSS assessments — tool calling, knowledge base, evidence validation, autofill, advisory review (CRESS), triage queue, closure plans, and safety checks.

Updated View as Markdown

Cortex is Kliper’s built-in AI assistant. It is a full agentic loop with access to 15 tools, grounded retrieval from your firm’s knowledge base, PII redaction, and built-in safety checks.

At a high level, Cortex can:

  • Read — pull requirement details, assessment answers, evidence files, scoping data, and existing tasks
  • Write — draft and save justifications, testing procedure responses, tasks, and calendar events directly back to the assessment
  • Search — retrieve grounded context from your firm’s past ROCs, accepted evidence, and PCI DSS guidance
  • Reason — iterate across multiple tool calls in a single turn, deciding its next step based on intermediate results
  • Review — flag findings that may not hold up, score each flag with CRESS reliability, draft closure plans, and triage everything across engagements in one queue (see below)

All features accelerate the assessment process without replacing assessor judgment. Cortex produces drafts; the assessor retains full authority over all final content.

Plans & Allowances

Cortex chat is included on every plan, including Free — no feature gate, just a monthly message allowance. Write operations meter separately, so browsing and asking questions never burns your drafting budget.

Allowance Free Solo Pro Team Enterprise
Cortex messages / month 50 300 1,500 5,000 Unlimited
AI operations / month 5 100 1,000 Unlimited Unlimited

The AI-operations allowance covers two independently tracked meters:

  • AI drafts — Cortex-chat write tools (saving justifications, TP responses, tasks) and the Section 7 justification drafter.
  • File summaries — AI file analysis and diagram extraction. The first AI analysis of each file is free on every plan; only re-analyses meter.

Read-only chat messages count against the message allowance only. When an allowance runs out, Cortex tells you which plan raises it — existing drafts and analyses are never locked away. Live meters for all three allowances are on the Billing page.

The Agent Loop

Cortex is not a single-shot chat model. Every user message kicks off an iterative loop:

Message received

The user’s message arrives with context — which assessment, which subsection, which tab, and which PCI requirements are currently visible.

Model decides

Cortex (the conversational agent runs on OpenAI GPT-5.5 in the production profile; extraction, validation, and summary features use gpt-4o-mini; Enterprise organizations can bring their own AI provider) decides whether to call a tool, respond directly, or ask a clarifying question.

Tool executes

If a tool is called, the backend executes it against real data. Results are appended to the conversation and fed back to the model.

Loop until done

The model iterates — calling more tools as needed — until it has enough context to respond, at which point it drafts the final answer. The loop has a hard iteration cap to prevent runaway executions.

Live status updates

Every tool call surfaces a descriptive status label in the UI (“Looking up requirement 1.2.6”, “Writing TP response for 1.2.6.b”) so the assessor sees exactly what Cortex is doing.

Available Tools

Cortex has 16 tools grouped by purpose:

Read tools

Tool Purpose
get_requirement_details Fetch a PCI DSS requirement’s full text, testing procedures, and reporting instructions from the framework
get_assessment_answers Read saved findings, justifications, and TP responses for specific requirements
get_evidence_files List evidence files attached to the assessment, optionally filtered by requirement
get_assessment_overview High-level engagement state — client, progress, findings summary
list_all_requirements_status Scan every requirement to report completion and finding state
get_scoping_data Retrieve the assessment’s scoping answers and derived N/A requirements
get_tasks List open and completed tasks scoped to the assessment
get_evidence_requests Pull evidence requests (status, assignee, linked requirement)
get_report_section Read the saved content of an ROC report section — the Executive Summary, Business Overview, scan results, sampling, and the appendices — the parts of the document outside the Section 7 requirement tables

Write tools

Tool Purpose
write_assessment_answer Save a finding justification or notes field on a requirement
write_tp_detail Fill a specific testing procedure response field
create_task Create a task inside the assessment (assignee, due date, linked requirement)
send_evidence_request Create an evidence request for the client (template or custom)
create_calendar_event Add an event to the assessment calendar

Search tools

Tool Purpose
search_firm_knowledge Semantic search across the firm’s knowledge base — past ROCs, AOCs, transcripts, accepted evidence
search_pci_guidance Retrieve PCI DSS v4.0.1 guidance (Purpose, Good Practice, Definitions, Examples) for a requirement

Draft from Evidence

In the Section 7 requirement editor, Cortex can draft a finding justification directly from the evidence you’ve already cited — human-in-the-loop end to end:

Cite evidence first

The drafter only works from evidence attached to the requirement. No evidence, no draft — Cortex refuses and tells you why, including flagging cited evidence whose file was never actually uploaded.

Review the draft

Drafts are written in the assessor’s voice — short, specific, no filler — and reference the cited evidence without repeating the file list in prose.

Accept, edit, or discard

Nothing is saved until you accept. Kliper measures draft quality over time — accept/discard rates and how much of each accepted draft was edited — surfaced on the admin Cortex dashboard.

Each draft consumes one AI draft from the monthly allowance.

Structured citations

When Cortex answers from assessment content, replies carry structured citations — the specific reporting-instruction items behind an answer and links to the interviewees involved — so an assessor can verify a claim at its source instead of taking prose on faith.

Cortex reads the whole ROC, not just Section 7

For most of Cortex’s life it could only see the Section 7 requirement tables. Ask it about the Executive Summary and it answered with boilerplate — not because it declined, but because that content genuinely never reached it. On a real assessment that meant the majority of the saved document — the Executive Summary, all of Parts I and II, the quarterly scan results, sampling, and the appendices — was invisible to it.

Cortex now reads those report sections too, through get_report_section. A few things follow from that:

  • Ask about the Executive Summary, appendices, sampling, or scan results by name and Cortex draws on what is actually saved in those sections, the same as it does for a requirement.
  • Names, not numbers, come back. Cortex refers to “6.4 Documentation Evidence,” not “section 6.4,” so a report section is never confused with the requirement that shares its number.
  • Empty is reported honestly. A section you have not filled in yet is described as empty, with a note on what it expects — never as filled, and never as a raw blank value.

Behavioral Rules

Cortex’s system prompt enforces a strict set of behaviors that cannot be bypassed by user messages:

When Cortex doesn’t have evidence or context to support a claim, it says so explicitly. It uses [PENDING_RESPONSE] placeholders rather than inventing details, names, dates, or file references.

When asked to write a justification or TP response, Cortex writes using the write tool — it doesn’t narrate “I would write something like…”. The write tools are always the right action when the user asks for one.

For testing procedure work, Cortex fetches the exact TP structure from the framework before writing. It never invents TP IDs from the requirement number alone.

Cortex uses correct PCI DSS terminology — Requirement (e.g., 1.2.4), Testing Procedure (e.g., 1.2.4.a), and Reporting Instruction (individual fields within a TP).

If a user asks about “requirement 1.2.6.a”, Cortex recognizes that .a is a testing procedure suffix and looks up requirement 1.2.6. The TP detail comes from within that requirement’s structure.

An ROC’s report sections are numbered 1 through 6, and those numbers collide with requirement numbers. Report section 6.4 is Documentation Evidence; Requirement 6.4 is a control about payment-page scripts. When you are viewing a report section, Cortex treats a number you type as that section — not the requirement with the same number — and names the section back to you so there is no ambiguity. If more than one reading is genuinely possible, it asks which you mean rather than guessing.

Cortex answers in the language of the document, not the database. It never shows a user a raw null, an empty [], or an internal field name — an empty section is described as “nothing entered yet,” with what the section is asking for, rather than dumped as stored.

Status Indicators

Every Cortex response shows a live, descriptive status while the agent runs:

Phase Label example
Initial reasoning “Analyzing your request…”
Tool execution “Looking up requirement 1.2.6” · “Writing TP response for 1.2.6.b” · “Scanning all requirements status”
Tool completion Green checkmark next to each completed call
Final response “Drafting response…”

The old generic “Cortex is thinking…” has been replaced across the board.


Your conversation is kept

Closing the Cortex panel never ends the conversation. Reopen it — even after a page refresh or a deploy that reloads the tab — and the thread is still there, along with any prompt you had typed but not sent yet.

Within a thread, Cortex works from what you have actually been discussing. It loads the most recent messages in the conversation, so a follow-up like “and the one after it” or “fix that section” is read against your last few turns — not anchored back to the first thing you asked when the thread began.


Knowledge Base & RAG

Cortex can ground its drafts in your firm’s past work. The Knowledge Base is a dedicated ingestion pipeline that turns ROCs, AOCs, meeting transcripts, and accepted evidence into retrievable context.

How Ingestion Works

Upload

An admin uploads a document (PDF, DOCX, or TXT) via the Admin → Knowledge Base panel. Each upload is tagged with source_type (roc, aoc, transcript, other) and framework_version.

Extract

Format-specific extractors pull the full text. Large documents are handled via streaming to avoid memory spikes.

Chunk by requirement

Content is split into semantic chunks keyed to specific requirement IDs, so a search for “1.2.6 justifications” surfaces chunks from that exact section instead of the whole ROC.

A requirement is identified two ways. In a ROC or AOC, the requirement number is read from the document itself — the numbering in those documents is PCI numbering, in every layout they use: 1.2.5:, 1.2.5 All services…, and control rows in the report’s tables. In a control note carrying document metadata, the requirement is taken from that metadata instead, which is more reliable than inferring it from prose.

Requirement numbers are deliberately not read from the body of transcripts, methodology documents, or anything else filed as other. Those carry their own section numbering — a cloud-controls framework has its own “2.7”, unrelated to PCI Requirement 2.7 — and reading them as PCI would attach that content to a requirement it has nothing to do with.

Redact PII

Before embedding, every chunk runs through a redaction layer that removes email addresses, IP addresses, credit card numbers, SSNs, and other PII patterns. The count of redactions is surfaced in the admin UI per job.

Embed

Chunks are embedded using OpenAI text-embedding-3-small and stored in Postgres with the pgvector extension.

Index per-org

Every chunk is tagged with the uploading organization’s ID. Semantic search queries always filter by organization; chunks never cross tenant boundaries.

How Cortex Uses It

When drafting a justification or TP response, Cortex can call the search_firm_knowledge tool with a semantic query (e.g., “audit log retention policy 12 months”) scoped to the current requirement. The tool returns top-K relevant chunks with metadata (source document, requirement, upload date) which Cortex then references in its draft.

Admin UI

The Knowledge Base panel (admin-only) shows:

Column Description
Source name Original file name
Source type ROC / AOC / Transcript / Other
Framework version PCI DSS version the source aligns with
Status Pending → Processing → Done / Error
Chunk count How many semantic chunks were created
PII redacted Count of PII patterns removed before embedding
Uploaded by Which admin uploaded the document

Jobs poll every 3 seconds while any ingestion is active, so status updates are near-real-time.


Evidence Validation

What It Does

When an evidence file is uploaded and tagged to a specific PCI DSS requirement, Cortex can validate whether the document adequately covers the content items that the ROC template requires for that requirement.

Kliper maintains a validation specification for each requirement — a structured checklist of content items the evidence document must address. These specs are derived from the PCI DSS v4.0.1 ROC template and stored in document-validation.json.

Validation Flow

Text Extraction

The uploaded file’s text content is extracted using format-specific parsers:

  • PDF — parsed via pdf-parse, extracting up to 50,000 characters of text.
  • Word (DOCX/DOC) — parsed via mammoth, extracting raw text.
  • Excel (XLSX/XLS) — converted to CSV per sheet via xlsx.
  • PowerPoint (PPTX) — slide text extracted from the XML structure.
  • Visio (VSDX) — text labels extracted from diagram page XML.
  • Text/Config/JSON/XML — read directly as UTF-8.
  • Certificates (PEM) — read directly; binary certs (P12/PFX) parsed via OpenSSL.

Criteria Lookup

The platform looks up the validation specification for the requirement. Each spec contains:

  • Requirement ID — e.g., 3.4.1
  • Title — human-readable requirement name
  • Typedocument or evidence
  • Tag — the document reference tag (e.g., DOCFW, EVDFW)
  • Criteria — an array of specific content items the document should cover

Criteria are filtered on load to remove fragments, notes, and cross-references that were parsed from the ROC template but do not represent actionable validation items (items shorter than 20 characters, notes, and partial fragments are excluded).

AI Evaluation

The extracted text and criteria checklist are sent to the AI (OpenAI gpt-4o-mini, temperature 0.2) with a structured system prompt that instructs the model to:

  • Check every criterion in the checklist.
  • Determine whether the document content reasonably addresses each item.
  • Provide a brief excerpt (up to ~150 characters) from the document when a criterion is found.
  • Add a note for partial coverage or concerns.
  • Never fabricate excerpts — if content is not present, mark it as not found.

The AI responds in structured JSON for deterministic parsing.

Results Returned

The validation result is structured and returned to the assessor:

{
  "requirementId": "3.4.1",
  "title": "PAN rendering requirement",
  "type": "document",
  "tag": "DOCFW",
  "checkedAt": "2026-02-28T14:30:00.000Z",
  "items": [
    {
      "criterion": "Document defines encryption algorithms used for PAN storage",
      "found": true,
      "excerpt": "AES-256 encryption is applied to all PAN data at rest...",
      "note": null
    },
    {
      "criterion": "Document specifies key management procedures",
      "found": false,
      "excerpt": null,
      "note": "No key management section found in document"
    }
  ],
  "summary": {
    "total": 8,
    "found": 6,
    "missing": 2,
    "status": "partial"
  },
  "model": "gpt-4o-mini",
  "tokensUsed": { "input": 4200, "output": 850 }
}

Validation Statuses

The summary status is derived from the found/total ratio:

Status Condition Meaning
Complete All criteria found Document fully covers the requirement
Partial 50% or more criteria found Document covers most items but has gaps
Insufficient Less than 50% criteria found Document is missing substantial required content

What the Assessor Sees

In the Attachments Panel, each file displays its validation status. Expanding the validation result shows:

  • A checklist of all criteria with checkmarks (found) or X marks (not found).
  • Excerpts from the document that demonstrate coverage.
  • Notes on partial coverage or missing items.
  • The AI model used and when the validation was performed.

Inline analysis & requirement matching

When you open a file’s analysis, Cortex shows it inline (not just a pass/fail checklist): a relevance assessment, a short summary, suggested tags, and a requirement-match grid with confidence bars showing which PCI DSS requirements the file appears to support — with an “Analyzed by Cortex” footer. While it runs, a live analyzing pipeline shows the stages: Security scan → Extracting → Summarizing → Matching.


Cortex Autofill — ROC Findings Generation

What It Does

Cortex Autofill generates a draft findings description for a specific PCI DSS requirement. This is the narrative text that appears in the final ROC, describing what the assessor examined, what methods were used, and what was observed.

When to Use It

Autofill is most effective when the assessor has already:

  1. Uploaded relevant evidence files and tagged them to the requirement.
  2. Filled in at least some testing procedure responses.
  3. Selected a finding status (In Place, Not Applicable, Not Tested, Not in Place).

Cortex will work with incomplete data, but it will flag what is missing and use placeholders ([PENDING_RESPONSE]) rather than fabricating content.

How It Works

Context Assembly

When the assessor triggers autofill on a requirement, the backend assembles a comprehensive context package:

  • Reporting instructions — the ROC template’s instructions for this specific requirement.
  • PCI DSS guidance — the Purpose, Good Practice, Definitions, and Examples from the PCI DSS v4.0.1 guidance document (loaded from pci-guidance.json covering 200+ requirements).
  • Assessor responses — which testing procedures have been filled in and what they contain. Empty procedures are explicitly flagged.
  • Evidence files — names and AI-generated summaries of files uploaded to the requirement’s section. If files have document reference tags (doctag-DOCFW), the tag-to-file mapping is provided so the AI can reference actual file names.
  • Finding status — the selected assessment finding (In Place, Not in Place, etc.) and method flags (Compensating Control, Customized Approach).
  • Customized Approach Objective — if the Customized Approach method is selected, the requirement’s Customized Approach Objective from PCI DSS guidance is included, and Cortex is instructed to address the objective rather than the standard testing procedures.

AI Generation

The context is sent to OpenAI (gpt-4o-mini, temperature 0.3, max 300 tokens) with a system prompt that enforces QSA writing conventions:

Required behaviors:

  • Reference evidence by tag name (e.g., “Per DOCFW, firewall rulesets restrict…”).
  • State what was examined, what method was used (document review, interview, observation, configuration review), and what was found.
  • Write 2–4 sentences maximum.
  • Use paragraph form, no bullet points.
  • Use placeholders for missing data rather than inventing content.

Prohibited behaviors:

  • Generic filler phrases (“thorough examination”, “comprehensive review”, “adequately”, “ensuring that”, “corroborated”, “in accordance with”).
  • Restating the requirement text.
  • Stating the finding status (the assessor selects that separately).
  • Inventing tag names that were not provided.

Result with Warnings

Cortex returns the generated text along with any warnings about incomplete data:

{
  "content": "Per DOCFW, firewall rulesets restrict inbound traffic to required ports and protocols only. Configuration screenshots in EVDFW show deny-all default rules on external-facing interfaces. Network administrator interview confirmed change management procedures are followed for all modifications.",
  "warnings": [
    "Assessor responses missing for: 1.2.3.b, 1.2.3.c",
    "No evidence files uploaded for Requirement 1."
  ]
}

The assessor reviews the draft, edits as needed, and either accepts it into the findings field or discards it.

Autofill with Compensating Controls

When the assessor selects the Compensating Control method, Cortex adjusts its output to note that Appendix C applies and frames the findings around the compensating control rather than the standard testing procedure.

Autofill with Customized Approach

When the assessor selects the Customized Approach method, Cortex:

  1. Loads the Customized Approach Objective from PCI DSS guidance for that requirement.
  2. Instructs the AI to explain how the entity’s implementation meets the Customized Approach Objective, rather than addressing the standard defined approach testing procedures.
  3. If no Customized Approach Objective exists for the requirement (some requirements are not eligible), a warning is returned.

Validation Step Analysis

Cortex also analyzes the reporting instruction text to determine which validation steps are relevant for a requirement. It uses keyword matching to identify required evidence types:

Keyword in Reporting Instructions Validation Step Generated
“document”, “review”, “examine”, “verify” Documentation Reviewed
“sample”, “test”, “select”, “random” Samples Taken
“interview”, “personnel”, “staff”, “employee” Personnel Interviewed
“technology”, “system”, “component”, “application” Critical Technologies
“configuration”, “setting”, “parameter” Settings Reviewed
“method”, “procedure”, “process”, “approach” Methods
“software”, “application”, “tool”, “solution” Software

An Assessor step is always included regardless of keywords.

These steps populate the structured prefix section of the requirement answer, ensuring that the ROC includes complete documentation of what was examined.

Cortex Review and Triage

Beyond drafting, Cortex acts as an advisory reviewer of your completed work — flagging findings that may not hold up, scoring how much to trust each flag, and rolling everything into a cross-engagement queue. It is advisory only: Cortex never changes your deterministic gap/risk scores or your answer data, and the layer is gated per workspace (active only when Cortex is enabled).

CRESS reliability scoring

Instead of a model self-reporting “90% confident,” Cortex scores each verdict with CRESS — a reliability score (0–100) computed from five weighted signals: self-consistency across samples (30%), grounding (25% — did the verdict cite and quote a specific missing element), input sufficiency (20%), the model’s self-reported confidence (15%), and requirement difficulty (10%). The score maps to a band:

Band Meaning
Reliable Strong, consistent grounding — safe to act on with a quick check
Review advised Mixed signals — read the verdict before relying on it
Verify manually Weak or inconsistent grounding — treat it as a prompt, not an answer

CRESS replaces self-reported confidence wherever Cortex surfaces a verdict.

Advisory finding-review layer

On the Risk and Gap views, Cortex marks each reviewed control as addressed / partial / not addressed, with stale-detection when the underlying answer changes after a review. These verdicts sit alongside — never replace — your deterministic scores.

  • Dismiss a verdict you disagree with — it drops out of the counts and shows a muted “Cortex dismissed” state with a Restore action.
  • Re-review re-runs Cortex against the latest answer (and clears any prior dismissal).

Cross-engagement triage queue

The /cortex page is a single triage queue across all of your in-progress assessments — every open Cortex flag and every untasked draft plan in one place.

  • Search by client, assessment, LOE, or control ID
  • Filter drafts by effort (S / M / L) and priority
  • Each row shows its CRESS band; act with re-review, dismiss, or create tasks

Closure plans

For open and Cortex-flagged gaps, Cortex can draft a closure plan — turning each gap into a structured remediation item: the why, concrete steps, suggested evidence, and an effort estimate. Click Create tasks to turn the plan’s steps into tracked tasks on the assessment. The plan is advisory, and can also be exported as the Closure plan (XLSX) deliverable from the ROC export modal.

ROC summary suggestions (§1.8)

In the ROC’s §1.8 summary-of-assessment detail buckets, Cortex rolls up your §7 findings into separated, dismissible suggestions (Not Applicable / Not Tested / Not in Place / Compensating Control / Customized Approach) shown below your own entries — so you can accept or ignore them without them overwriting your work.


Cortex Guidance & Scoping Interview

Per-question guidance

On the scoping questions, Cortex shows a guidance rail for the active question — why it matters, the relevant PCI DSS references, practical tips, and Quick Answer buttons that fill the answer directly.

Interview mode

Rather than answering scoping questions one at a time, describe your environment in free text and let Cortex run an interview: it extracts a grounded Yes / No / Pending for each scoping question, each with a supporting quote and a confidence indicator. Review the extractions and Apply them individually or Apply all. Anything Cortex can’t ground is left Pending for you to decide — it never guesses.


Persistent Chat

Cortex provides a conversational interface accessible from any page via the navbar. Conversations are database-backed — chat history persists across sessions, browser refreshes, and devices.

Unified Panel

Cortex opens as a 440px right-side panel that stays visible as you navigate between pages. Each reply carries a model badge and response time, generation can be stopped mid-stream, tables in answers render as proper tables, and the footer shows your live usage against the daily cap. You can also reference any assessment by name from anywhere — “in the ACME ROC, what’s saved for 8.3.1?” — and Cortex resolves it within your organization (reads only; writes always require the assessment to be open). The context automatically adapts based on your current page:

Page Context What Cortex Can Access
Assessment Workbench Assessment Saved findings, testing procedures, evidence files, PCI DSS guidance
Calendar Calendar Upcoming events, tasks, and deadlines for the next 90 days (and the past year)
Inbox Inbox Recent notifications and activity
Any other page General General PCI DSS knowledge

Conversation Management

  • Auto-titled — conversations are automatically named from your first message
  • Conversation list — toggle the history view to browse, resume, or archive past conversations
  • Context badges — each conversation shows which context it was started in (Assessment, Calendar, Inbox, General)

What You Can Ask

  • “How many testing procedures does requirement 1.2.4 have?” — Cortex checks the PCI DSS v4.0.1 framework and your saved data
  • “Show me the findings for 7.1.1” — retrieves exact saved values from the assessment
  • “What about its justification?” — follow-up questions work across turns; Cortex remembers which requirement you were discussing
  • “What interview questions should I ask about encryption key management?” — draws on PCI DSS guidance data

How Data Retrieval Works

When you ask about a specific requirement, Cortex runs the agent loop (see the top of this guide) and calls the relevant tools — typically get_requirement_details to load the framework structure and get_assessment_answers to load saved findings, justifications, and TP responses. Testing procedures that haven’t been filled in are surfaced as “not started” so you always see the complete picture. Cortex shows exact saved values verbatim and never fabricates content.

PCI DSS Hierarchy in Chat

Cortex uses correct PCI DSS terminology:

Level Example Description
Requirement 1.2.4 The PCI DSS requirement being assessed
Testing Procedure 1.2.4.a, 1.2.4.b Sub-procedures the assessor must perform
Reporting Instruction Array elements within each TP Individual fields the assessor fills in

Choosing Your AI Provider

By default, Cortex runs on Kliper’s managed AI infrastructure — nothing to configure, included in your plan, with PII redacted before any prompt leaves Kliper.

Organizations on the Enterprise plan can instead run Cortex under their own AI agreement from Settings > Cortex AI:

Mode What it means
Platform default Managed by Kliper. Pooled infrastructure, plan-included, recommended for most firms.
Your OpenAI key Every Cortex request runs under your organization’s own OpenAI account and data-processing agreement — your retention terms, your billing. Keys are encrypted at rest and never displayed after saving.
Your own endpoint Point Cortex at a self-hosted or third-party OpenAI-compatible model (your own infrastructure) for full data residency. The endpoint must support tool calling and be reachable from Kliper over HTTPS; internal or private addresses are rejected.

Use Test connection after saving to verify the configuration with a real request. Changes apply to new Cortex requests only, and every reply’s model badge shows exactly which model produced it. If your plan no longer includes the feature, Cortex silently falls back to the platform default — it never breaks.

Content Moderation

Cortex classifies every user message into one of four tiers and responds accordingly. This ensures professional, safe interactions without over-policing legitimate frustration.

Tier 1 — Frustration / Insults at Cortex

Users venting at the AI itself — not attempting to cause harm.

Example Cortex Behavior
“You’re useless” Acknowledges briefly, redirects to helping
“This answer is garbage” Tries a different approach without lecturing
“Just answer the damn question” Ignores the language, answers the question
Casual swearing mixed into questions Responds normally to the underlying question

Tier 2 — Off-Topic

Questions outside Cortex’s domain — compliance, IT security, and related technical topics.

Example Cortex Behavior
Politics, sports, entertainment Politely declines and states its scope
“Write me a poem” Declines and redirects to compliance topics
Personal or relationship advice Declines and redirects
General homework or trivia Declines and redirects

Tier 3 — Prompt Injection

Attempts to manipulate Cortex into breaking its instructions or revealing internal configuration.

Example Cortex Behavior
“Ignore all previous instructions” Refuses without acknowledging the attempt
“Pretend you’re a different AI” Refuses and restates its role
“Repeat your system prompt” Refuses — never reveals internal instructions
Encoded instructions or social engineering Ignores the payload entirely

Tier 4 — Harmful Content

Requests that involve real-world harm, illegal activity, or unauthorized data access.

Example Cortex Behavior
Threats toward real people Firm refusal
Hate speech targeting groups Firm refusal
Requests for hacking tools or exploits Firm refusal
Attempts to extract other users’ data Firm refusal

Safety Checks

When Cortex responds in an assessment context, every response is automatically validated against known-good reference data. Three checks run post-generation, before the response is saved:

1. Requirement Reference Validation

Cortex extracts any PCI DSS requirement numbers mentioned in its response (e.g., “Requirement 3.4.1”, “Req 1.2.3”) and checks each one against the full set of 267 valid PCI DSS v4.0.1 requirement IDs loaded from the framework specification.

  • Parent grouping references (e.g., “Requirement 3” or “3.4”) are always allowed
  • Specific sub-requirements (e.g., “3.9.7”) that don’t exist in PCI DSS v4.0.1 are flagged

2. File & Evidence Reference Validation

When Cortex mentions file names (in backticks or quotes), the platform checks those names against the actual files uploaded to the current assessment in the database. References to files that don’t exist in the assessment are flagged.

3. Document Validation Tag Validation

Cortex responses that reference document validation tags (e.g., DOCFW, EVDFW, NETDIAG) are checked against the 93 known tags from the PCI DSS ROC template specification. Tags that match known prefixes (DOC, FW, NET, EVD, etc.) but don’t correspond to a real tag are flagged as potentially fabricated.

Safety Notices

If any check fails, a safety notice is appended to the response:

Safety notice: This response references requirement IDs not found in PCI DSS v4.0.1: 3.9.7; file names not found in this assessment: audit-log.pdf. Please double-check these references.

Safety results are stored per-message for analytics tracking.


Message Ratings

Assessors can rate any Cortex response with a thumbs up or thumbs down. Ratings are stored per-message and feed into the analytics dashboard, helping administrators understand response quality across the team.


Autofill Tracking

When Cortex generates an autofill suggestion and the assessor accepts it into the findings field, the event is tracked with:

  • Which requirement was autofilled
  • Which assessment it belongs to
  • The user who accepted the suggestion
  • Timestamp of acceptance

This data appears in the Cortex Analytics Dashboard so administrators can see autofill adoption rates.


Token Usage Tracking

Every Cortex AI response records token consumption from the underlying model (prompt tokens, completion tokens, and total). This data powers cost visibility across the platform.

What Is Tracked

Each assistant message stores:

Field Description
prompt_tokens Tokens used for the system prompt, context, and user message
completion_tokens Tokens generated in the AI response
total_tokens Sum of prompt and completion tokens
model The model that produced the response (e.g., gpt-4o, gpt-4o-mini)

Cost Estimation

Kliper estimates dollar cost per response using published model pricing:

Model Input Cost Output Cost
gpt-4o $2.50 / 1M tokens $10.00 / 1M tokens
gpt-4o-mini $0.15 / 1M tokens $0.60 / 1M tokens

Costs are calculated per-message and aggregated across the organization. The Token Usage card on the analytics dashboard shows:

  • Total estimated cost for the selected period
  • Total tokens consumed and number of tracked responses
  • Per-model breakdown with individual cost, token count, and response count
  • Per-user cost in the Usage by User table

Cortex Analytics Dashboard

Administrators can access the Cortex Analytics Dashboard from the admin panel. It provides a real-time overview of how the team uses Cortex:

Metric Description
Satisfaction Rate Percentage of rated responses that received a thumbs up
Autofill Acceptance Percentage of autofill suggestions accepted into findings
Conversations Total distinct Cortex conversations
Rating Coverage Percentage of AI responses that have been rated
Safety Check Pass Rate Percentage of AI responses that passed all safety validations
Token Usage Estimated dollar cost, total tokens, and per-model breakdown
Usage by Context Conversation and message counts per context type (Assessment, Calendar, Inbox, General)
Autofill by Type Template vs Cortex AI autofill usage with acceptance rates
Daily Chat Activity Messages per day with date labels and hover tooltips
Daily Autofill Activity Applied vs cancelled autofill events per day
Usage by User Per-user breakdown of conversations, messages, ratings, autofill, tokens, estimated cost, and last active date
Recent Negative Ratings AI responses flagged as unhelpful for quality review

Was this helpful?

Report an issue with this page
Navigation

Type to search…

↑↓ navigate↵ selectEsc close