What Evaluation Panels Look For in Each Criterion

Most government tender criteria fall into five families. Here's what a panel is assessing in each one — and the mistake that costs points every time.

Lewis Heard, Founder — ProcureHQ
10 min read

Written by Lewis Heard, founder of ProcureHQ. Circa 15 years as a Government Procurement Manager. 1,000+ tender responses personally received and reviewed, and 500+ evaluation meetings chaired — scoring done by the evaluation committee he chaired, across procurements totalling over $1 billion in combined contract award sums.

Evaluation criteria look different on every tender. The wording changes, the order changes, the number of them changes. But after a few hundred submissions you notice that almost every non-price criterion in Australian government procurement belongs to one of five families.

Knowing which family a criterion belongs to tells you what the panel is actually trying to establish — which is rarely what the heading says on its face.

Family 1: Organisational capability and capacity

Typical wording: "Relevant experience of the tenderer", "Demonstrated capability to deliver works of this nature and scale", "Organisational capacity".

What the panel is establishing: two separate things, and contractors routinely answer only the first.

Capability is whether you have done this kind of work. Capacity is whether you can do it now, alongside everything else you are currently contracted to deliver. A panel reading a strong experience section with no mention of current workload is left with an unanswered question, and unanswered questions do not score well.

What moves the score:

  • Comparable projects, not impressive ones. A $40m job that shares nothing with this one is weaker evidence than a $3m job on an occupied site with the same access constraints. Comparability beats scale.
  • Named projects with client, value, dates, and the specific feature that makes it comparable.
  • A clear statement of current commitments and available resource. Financial capacity often belongs here too, even where the criteria are silent on it — on head contracts it is standard for a panel to want it.

The mistake: the capability statement that lists twenty projects with no explanation of why any of them are relevant. Twenty unexplained projects is a weaker answer than three explained ones.

Decision Signal 1 of circa 150 — Comparable Project Experience. Panels do not reward volume of experience; they reward evidence that you have solved this problem before. The closer your nominated projects sit to the scope, scale, site conditions and delivery environment of the tender, the more confidently a panel can score you.

Family 2: Personnel

Typical wording: "Proposed project structure and key personnel", "Capability and experience of the project team".

What the panel is establishing: whether the people named will actually be on this job, and whether the structure makes sense for the work.

Panels are experienced enough to know that CVs are recycled. What they look for is the join between the person and the project: this superintendent, on this site, doing this role, with this much of their time.

What moves the score:

  • An organisation chart that shows this project's structure, including the interfaces to the client and to subcontractors — not a generic corporate hierarchy.
  • CVs tailored to the role being proposed, leading with the projects that resemble this one rather than a career chronology.
  • Committed time allocations. "Full-time on site for the duration" scores; silence on availability invites the assumption that the person is spread across four jobs.
  • Named succession or backfill arrangements. Panels have been burned by key personnel disappearing after award, and a response that acknowledges the risk reads as more credible, not less.

The mistake: a twelve-page CV pack for people whose role on this project is never stated.

Family 3: Technical capability — methodology and program

Typical wording: "Methodology and tender program", "Proposed approach to the works", "Construction methodology".

This is usually the most heavily weighted non-price criterion, and it is where the widest scoring spread appears in moderation. It is the criterion most worth your time.

What the panel is establishing: whether you understand this job, and whether your plan for it is credible.

What moves the score:

  • Site-specific constraints named and addressed. Adjacent operations, restricted hours, live services, marine access windows, heritage fabric, occupied buildings — whatever the tender documents describe, your methodology should show you read them.
  • A sequence, not a list of activities. Panels are reading for logic: why this order, what depends on what, what happens when the critical item slips.
  • A program that ties to the methodology. When the narrative describes a staging approach the program does not reflect, both scores fall, because the panel now doubts the plan rather than just the document.
  • Risks specific to this job, with the control you will actually apply. A generic risk register with "adverse weather — monitor forecast" contributes nothing.

The mistake: the methodology that describes how your company builds things in general. It is grammatically fine, professionally presented, and reads as flat as paper in a room where the assessor has just read a competitor who named the site's access constraint by name.

Decision Signal 2 of circa 150 — Project-Specific Methodology. Generic methodology is the single most common cause of a mid-band score on the highest-weighted criterion. Panels distinguish between a contractor describing their standard process and a contractor describing how they will deliver this scope, on this site, under these constraints.

Family 4: Management systems and plans

Typical wording: "Management systems and plans", "Quality, safety and environmental management", "Degree of compliance with the proposed contract".

What the panel is establishing: whether your systems are real and whether they will be applied to this project.

Almost every tenderer holds certification. Certification alone is therefore a differentiator of essentially zero value — everybody clears the bar, so nobody gains ground on it.

What moves the score:

  • Project-specific application. Not "we hold ISO 9001", but what the inspection and test plan looks like for the key work of this scope, who signs off hold points, and how non-conformances get closed.
  • Extracts and evidence rather than whole system manuals. A panel does not want your 90-page quality manual; it wants the two pages that show the system working on a job like this one.
  • Named systems, named documents, named responsibilities. "The site manager will maintain the ITP register" is scoreable. "Quality is embedded in everything we do" is not.
  • Where the criterion covers compliance with the proposed contract, deal with departures honestly and early. A clean, clearly stated position reads better than a buried qualification the panel discovers on page 60.

The mistake: attaching certificates and system manuals in place of an answer. Attachments are evidence for a claim; they are not the claim.

Family 5: Social procurement and sustainability

Typical wording: "Local content and jobs", "Skills, training and diversity", "Aboriginal participation", "Sustainability outcomes", "Social procurement commitments".

What the panel is establishing: whether your commitments are specific, measurable and deliverable — because these commitments frequently become contractual, reported and audited.

This family has grown fastest in Australian government procurement and is where the widest gap sits between what tenderers write and what agencies want.

What moves the score:

  • Numbers with a basis. A stated number of apprentice hours or local spend percentage, with a short explanation of how it was derived from this project's scope, beats an enthusiastic paragraph of intent every time.
  • Named partners and suppliers where you have them, rather than an undertaking to identify some later.
  • Mechanisms for delivery and reporting: who is responsible, how it is tracked, how it is reported to the agency.
  • Honesty about what is achievable. Panels have seen commitments that could not survive contact with the program. An overreaching number invites scepticism about the rest of your submission.

The mistake: treating this as a values statement. It is a deliverables criterion wearing a values heading.

The thread running through all five

Across every family, the same distinction separates a mid-band score from a high one.

Decision Signal 3 of circa 150 — Evidence Over Claims. A claim tells a panel what you assert about yourself. Evidence lets a panel verify it. Because consensus scores must be defended in a moderation meeting, evidence is what a panel member uses to argue your score upward — and a claim gives them nothing to hold.

Read your draft with one question: for each assertion, what would the assessor point at if a colleague challenged the score?

If the answer is a named project, a document extract, a drawing, a number with a derivation, or a specific commitment — that assertion is doing work. If the answer is "the paragraph says so", it is not.

Seeing it on your own submission

The five families above are how ProcureHQ organises the assessment. The platform extracts the evaluation criteria from your actual tender documents, maps each one to the family it belongs to, and assesses your response against the published criteria and their assessment requirements — returning a score per criterion, the findings behind it, the evidence relied on, and what would lift it.

If you have not written the response yet, the Smart Submission Blueprint works the other way around: it builds the section structure for this specific tender, criterion by criterion, with what each section needs to contain and what evidence a panel expects to find there. It does not write your content — the words have to be yours, and a panel can tell when they are not — but you never start from a blank page.

Your first tender is free, one per organisation: a full evaluation and a complete Smart Submission Blueprint, at no cost.


Frequently asked questions

How many evaluation criteria does a typical government tender have? Commonly four to six non-price criteria, each often carrying stated assessment requirements underneath. Some tenders publish a weighting against each criterion and some do not — where weightings are published, they tell you where to spend your effort.

Should I structure my response around the evaluation criteria? Yes. Mirroring the tender's criteria and their stated requirements, in the tender's own order, makes the panel's job straightforward. Reorganising the structure into your own preferred order forces every assessor to do translation work while scoring you.

Does holding ISO certification improve my score? Rarely on its own, because most tenderers hold it. What scores is showing how the certified system will be applied to this project — inspection and test plans, hold points, responsibilities and reporting.

What is the most heavily weighted criterion usually? Methodology and program is frequently the highest-weighted non-price criterion, and it produces the widest spread of scores in moderation — which makes it the highest-return place to invest your writing time.

Related insights