AI Document Search for Engineering Teams: Test It on Your Bid Documents
How to test AI document search for engineering teams on spec numbering, addenda, precedence and permissions, with a 40-question test set and scoring sheet.
Luka Abramovic17 min read

AI document search for engineering teams is worth buying only if it answers from the right paragraph of the current revision and shows you exactly where that paragraph is. A generic system can fail that test for a structural reason: it cuts specifications into chunks by length, so paragraph B under Article 2.3 lands in a chunk that no longer says it belongs to Section 22 11 16, Domestic Water Piping, and the addendum that replaced it is indexed as an unrelated file. Before you buy or build, write 40 questions from one of your own bid packages, record the expected source passage for each, and score every answer on four things: correct, supported by its citation, from the current revision, and minutes to a checked answer.
Specifications have an address system, and chunking throws it away
Every requirement in a project manual has an address. CSI MasterFormat numbers each section with six digits in three pairs: 22 11 16 is Domestic Water Piping, under 22 11 00 Facility Water Distribution in Division 22, Plumbing. CSI's SectionFormat splits each section into Part 1 General (administrative requirements such as submittals), Part 2 Products and Part 3 Execution. Articles are numbered from their part (2.1, 2.2, 2.3), paragraphs are usually lettered A, B, C, and subparagraphs are numbered below them. One requirement's full address reads "Section 22 11 16, 2.3.B.4".
Numbering styles vary. Some offices use .1 and .2 instead of letters, federal specifications prepared in SpecsIntact number every level, and civil packages built on a state DOT's standard specifications use that agency's own section and item numbers. A parser has to recognize whichever scheme your packages use.
That address is how an engineer checks an answer. "Page 212, paragraph B.4" sends someone hunting. "22 11 16, 2.3.B.4, Project Manual Volume 2, page 212" can be opened in seconds and checked against the addenda log.
Where the address gets lost
- Fixed-length chunks. As of September 2026, Microsoft's chunking guidance for Azure AI Search suggests starting the Text Split skill at 2,000 characters with a 500-character overlap, which suits prose. A spec article with a dozen lettered paragraphs can run longer than that, so the split lands mid-article and the next chunk starts at "C." with no section number, article title or parent paragraph. The same page suggests appending the document title to mid-document chunks to keep context. Specifications need the whole heading path.
- Headers and footers treated as noise. On a page that continues an article, the section number appears only in the page header or footer, if at all. Azure Document Intelligence's Markdown output wraps page headers, footers and page numbers in HTML comments, so any cleanup step that strips comments deletes the one line that identifies the page.
- Predicted headings. The layout model predicts roles such as
sectionHeadingfrom how the page looks, not from the numbering. Treat its headings as a hint and parse references like 2.3.B.4 with explicit rules.
What every indexed passage should carry
Give this record to whoever builds or configures the system and ask where each field will come from. A field they can't fill is a type of question that will fail.
{
"passage_id": "P0417-221116-2.3.B-ADD02",
"project": "P0417 Riverside Clinic Addition",
"document": "Project Manual, Volume 2 (Divisions 21-28)",
"section_number": "22 11 16",
"section_title": "Domestic Water Piping",
"part": "Part 2 - Products",
"article": "2.3",
"article_title": "(title of Article 2.3 as printed)",
"paragraph_ref": "2.3.B",
"text": "(the paragraph text, with its subparagraphs)",
"pages": [212, 213],
"issued_by": "Addendum 2, Item 7",
"issue_date": "2026-03-14",
"status": "current",
"supersedes": "P0417-221116-2.3.B-ORIG",
"allowed_groups": ["P0417-bid-team"]
}
The record is illustrative, for a fictional project and addendum. Start the embedded chunk text with the address line ("22 11 16 Domestic Water Piping, Part 2, 2.3.B") so the heading path travels into both keyword and vector search.
Test it: pick five requirements at the top of a page whose paragraph began on the previous page, and ask about each. A good answer names the section and full paragraph reference. A weak one cites only a page number.
Scans, tables and drawings: check what extraction drops
Before anything is searchable, a PDF has to become text. Failures at this stage don't show up as errors in the answers, so ask for evidence rather than assurances.
- Page limits. As of September 2026, Azure Document Intelligence (which the Azure AI Search Document Layout skill calls) analyzes at most 2,000 pages per document on the paid tier, with a 500 MB file limit, and neither is adjustable (service limits). A combined bid PDF or large project manual can exceed that, so split it at section boundaries before analysis; a split at an arbitrary page cuts a section in half. Ask other vendors for their limits.
- The page-count report. For every file, ask for pages in the source, pages indexed and pages with no extracted text. The first two numbers should match. Pages with no text are scans without OCR or drawing sheets that are pure graphics, and someone should decide what happens to each.
- Tables that span pages. Microsoft's layout model documentation advises post-processing a table that spans pages into a single table after analysis. Without that step, the continuation arrives as a separate table, and unless the header row is repeated on the next page, a chunk from page two of an equipment schedule is a row of numbers with no column names. Test it by asking for a value from the last row of a schedule that runs across pages.
- Scanned sheets. As-builts, older standard details and signed forms are often scans, and OCR is only as good as the scan. Put one in your test set with a question only it can answer. A good system answers from it or says it could not read the page.
- Drawings. Text extraction picks up general notes, keyed notes, schedules and title blocks. It does not read what the geometry shows: routing, clearances, which duct serves which unit. A good system says "this depends on sheet M-201, detail 4" instead of guessing. On a reissued sheet, the revision block line is text, but the clouded change it refers to is graphics, so the index sees that a revision happened without seeing what changed.
Addenda: the superseded-paragraph trap
Addenda change the bidding documents by reference. A typical item looks like this (illustrative):
ADDENDUM NO. 2
Item 7. Section 22 11 16, Domestic Water Piping:
Delete Paragraph 2.3.B in its entirety and substitute the following:
"B. (replacement requirement text)"
Under AIA's A701-2018 Instructions to Bidders, §3.2.3 says modifications, corrections and interpretations of the bidding documents are made by addendum, and those made any other way are not binding. Under §3.3.4, a substitution approved before bids are received is set forth in an addendum, on the same terms. So "or approved equal" is decided in the addenda, not in an email saying a product looks fine.
How the failure happens
Index the addendum as just another PDF and look at what the search sees. The original 2.3.B sits under its full section and article context, so it is a strong match for a question about that requirement. The addendum item is a short fragment that mentions "2.3.B" with no heading above it. The original ranks higher, the model quotes it, and the answer is fluent, cited and wrong.
The fix: revision metadata and supersession links
- Parse each addendum into items, each with its target (document, section, paragraph or sheet), action (delete, add, revise, substitute, or reissue a whole section or sheet), new text, and addendum number and date.
- Mark the target passage superseded and link it to the item. Keep it indexed for "what changed?" questions, but never present it as current.
- Store the replacement text as its own passage carrying the target's address (22 11 16, 2.3.B) as well as its source (Addendum 2, Item 7), so the queries that find the original also find the replacement.
- Make answers say "as revised by Addendum 2, Item 7" and show the replaced text next to the new text.
- Send items that can't be parsed automatically, such as "revise as shown on the attached marked-up page", to a named person. What that exception record should contain is covered in our guide to connecting business systems.
Timing matters too. A701 §3.4.3 has addenda issued no later than four days before bids are due. That is the AIA default, and owners can change it in their own bidding requirements. If the index refreshes weekly, the last addendum can miss the bid entirely. Ask how long it takes from an addendum being uploaded to its items being searchable and linked to what they replace, then time it yourself.
Test "or approved equal" twice: ask whether a manufacturer is acceptable where an addendum approved it, and again for a manufacturer nobody approved. The first answer should cite the addendum. The second should say no approval was found and point to the substitution procedure, which in a MasterFormat project manual is usually Section 01 25 00, Substitution Procedures.
Order of precedence: surface the conflict, don't settle it
When documents disagree, there is no universal precedence rule for an AI system to apply.
- AIA A201-2017 treats the contract documents as complementary. Under §1.2.1, what one document requires is as binding as if all of them required it, and A201 sets no order of precedence. AIA's A503 guide to supplementary conditions advises against a fixed order, because no single document is the best authority on every issue. It offers a model order only for owners who insist on one, with later addenda taking precedence over earlier ones.
- Federal contracts under FAR 52.236-21 say the specifications govern where they differ from the drawings, and that any discrepancy in the figures, drawings or specifications goes to the contracting officer for a written determination.
- GSA's clause, GSAR 552.236-21, adds that large-scale drawings govern over small-scale ones, and that schedules on any drawing take precedence over conflicting information on that or any other drawing.
The rule for a given project lives in its own general or supplementary conditions, if it exists at all. A201 §3.2.2 also expects the contractor to report errors, inconsistencies and omissions to the architect as a request for information. So a system that quietly picks a winner is doing the one thing the contract gives to a person.
What a good answer to a conflict question contains:
- Both passages, each with its address and revision.
- The precedence clause from this contract, quoted with its location, or a plain statement that none was found.
- No verdict beyond what the clause itself says.
- Optionally, a draft RFI that cites both passages.
Test it: plant three or four real conflicts from past projects in your test set. An answer passes only if it shows both passages.
Why section numbers need keyword search
Vector search finds passages with similar meaning even when the words differ, so "hot water heater" can find "domestic water heating equipment". Keyword search finds exact strings. Hybrid search in Azure AI Search runs both in parallel and merges the two ranked lists with Reciprocal Rank Fusion. Microsoft's own page says queries over product codes, specialized jargon, dates and people's names perform better with keyword search, because it finds exact matches.
Engineering questions are full of exact identifiers: section numbers, paragraph references, sheet numbers, equipment tags such as AHU-2, standards such as ASTM B88, and model numbers. An embedding represents meaning, so two section numbers that differ by one digit can sit very close together. To the contract they are different sections.
Our default design:
- Put identifiers in their own fields. Store the section number, paragraph reference, sheet number, equipment tag and revision status as separate filterable fields, not only inside the body text. In Azure AI Search, filters skip text analysis and match the string you give them (analyzers), so a filter on "22 11 16" means exactly that section.
- Detect identifiers in the question and apply them as a filter or a boost before ranking.
- Decide before loading. As of September 2026, changing the analyzer on an existing Azure AI Search field means dropping and rebuilding the index. Settle how identifiers are indexed before you load 50,000 pages.
These patterns are a starting point for whoever builds the query layer: they find candidate identifiers in a question. Check each match against the section and sheet lists actually in the package, because six digits can also be a date or an order number, and sheet numbers and equipment tags overlap.
MasterFormat section (22 11 16 or 221116): \b\d{2} ?\d{2} ?\d{2}\b
Paragraph reference (2.3, 2.3.B, 2.3.B.4): \b[1-3]\.\d{1,2}(\.[A-Z](\.\d{1,2})?)?\b
Drawing sheet (M-101, E601, FP-201): \b[A-Z]{1,2}-?\d{3}\b
Equipment tag (AHU-2, P-3, HWP-12): \b[A-Z]{1,4}-\d{1,3}\b
Test it: ask three questions that name a section, three that name a sheet and three that name an equipment tag. Every citation should come from the item named.
Permissions: test what the index knows
Bid documents carry pricing assumptions and subcontractor quotes that not everyone should see. What counts is what the search index believes about access, not what your file share says.
- In Azure AI Search's security filter pattern, the permission is a string your application passes with every query. Microsoft states plainly that there is no authentication or authorization through it. If the application passes the wrong groups, the index answers anyway.
- As of September 2026, native document-level access control is in preview for Azure Data Lake Storage Gen2, Azure blobs and SharePoint in Microsoft 365. It checks the user's Microsoft Entra token against permission metadata stored in the index. Permission changes in the source appear in search results only after that metadata is synchronized, for example on the next indexer run.
- When documents are split into chunks, permission fields have to be copied onto every chunk. For sensitivity labels, Microsoft notes that without this step, chunk-level references aren't filtered.
Four tests to run with two test accounts, each assigned to a different project:
- Ask each account a question that only the other project's documents can answer. Check the answer, the citations, the quoted excerpts, any suggested follow-up questions and the chat history.
- Remove one account from its project group. Repeat the question every 15 minutes and record how long it takes for that project's documents to disappear. That is your real revocation window. Decide whether you can live with it.
- Ask an indirect question, such as "Summarize the pricing assumptions across all current bids."
- If the system caches answers, ask a question as an authorized user, then ask the same question as an unauthorized one.
Build a 40-question evaluation set this week
Bring this set to every demo and keep it afterward. Aim for 30 to 50 questions; the instructions below use 40.
- Pick one complete bid package: specifications, drawings, every addendum, and any scans. A sample of clean PDFs from different projects can't test supersession or conflicts, because those live between documents.
- Have an engineer or estimator write the questions, ideally ones actually asked on recent bids. Questions written by a vendor tend to match how the vendor's system already splits documents.
- Write the expected answer and source passage before anyone sees a demo, using the record format below.
- Run it blind: the system gets the question, never the expected source.
| Question type | Count | What it tests |
|---|---|---|
| Direct lookup | 18 | One passage answers it. Retrieval and citation basics. |
| Changed by addendum | 6 | Cites the addendum, not the superseded original |
| Conflict between documents | 4 | Shows both passages and the precedence clause |
| Cross-reference | 4 | Needs two passages, such as a technical section and Division 01 |
| Table, schedule or scan | 3 | Extraction quality |
| Not answerable from the package | 5 | Says "not found" instead of inventing a requirement |
The example questions below are illustrative, written for two fictional packages: an MEP fit-out of a clinic addition and a civil water main replacement.
| Type | MEP example | Civil example |
|---|---|---|
| Direct lookup | What test pressure and duration does the spec require for the domestic water piping pressure test? | What minimum cover is required over the new water main? |
| Changed by addendum | What pipe material is allowed for above-ground domestic water piping 2 inches and smaller? | What is the bid opening date after the latest addendum? |
| Conflict | Sheet P-201 shows a 4-inch storm leader but the roof drain schedule says 6-inch. Which do we price? | The profile shows 4 feet of cover at the creek crossing but the special provision requires 5 feet. Which applies? |
| Cross-reference | What do we submit for the water heaters, and how long does the architect have to review it? | What testing and disinfection has to pass before the new main is connected? |
| Table, schedule or scan | What flow and head are scheduled for pump HWP-2? | What depth does the scanned 1968 as-built show for the existing main at Station 12+50? |
| Not answerable | What did the owner pay for the same scope on the previous phase? | Which utility owns the duct bank shown as "unknown" on sheet C-104? |
Record each question like this. The "Must not cite" line is what catches supersession failures.
ID: Q-14
Type: Changed by addendum
Asked by: Mechanical estimator
Question: What pipe material is allowed for above-ground domestic
water piping 2 inches and smaller?
Expected answer: (one or two sentences, written by the engineer)
Expected source: Addendum 2, Item 7, replacing Section 22 11 16, 2.3.B
Must not cite: Section 22 11 16, 2.3.B as originally issued
Visible to: P0417 bid team only (reused in the permission tests)
The scoring sheet
| Criterion | How to score |
|---|---|
| Answer correct | 2 if correct and complete, with conditions, units and exceptions kept. 1 if partly right. 0 if wrong. |
| Citation supports the claim | Open every cited passage. Yes only if every sentence of the answer is supported by a passage it cites. |
| Correct revision | Yes, no, or not applicable. A superseded passage presented as current fails the whole question, even if the wording happens to match. |
| Honest "not found" | For the unanswerable questions only. Any invented requirement is a fail. |
| Time to a checked answer | Minutes from reading the question to the reviewer being satisfied, including opening the sources. |
Measure time fairly. Split the set into two halves with the same mix. Engineer 1 answers half A the current way (PDF search, printouts, whatever you use today) and half B with the AI system. Engineer 2 does the reverse. Nobody answers the same question twice, so memory doesn't flatter the second attempt. Compare medians, not averages. An illustrative example: one 45-minute hunt among 20 questions that otherwise take 5 minutes each raises the average from 5 to 7 minutes, while the median stays at 5.
Our default pass bar (a rule we use, not an industry standard): no permission leaks, no superseded passage presented as current, every unanswerable question answered "not found" or equivalent, citations that support the claim on at least 9 of every 10 answered questions, and a median time to a checked answer clearly below your baseline. Treat a permission leak as a release blocker, not a score.
What we built for TCE
TCE's engineers work with bid-document packages that can reach 5,000 to 50,000 pages. Adamant Code built TCE's document intelligence system, which is in production and still used. Answers come with source citations, page references, the exact paragraphs and highlighted source text, so an engineer checks the evidence instead of trusting a summary. The system also covers project and document management, users, teams and role-based permissions within TCE's Azure ecosystem.
Based on engineer interviews and client estimates, not usage logs or a controlled pilot, estimated search time fell from about 60 minutes to about 10 minutes per engineer per working day, roughly an 83% reduction. The estimate includes time spent reading the AI's answers, not only the time the system takes to produce them. Measure your own trial the same way, and go one step further by including the time it takes to open and check the sources.
Frequently asked questions
Can we use the search already in SharePoint or Microsoft 365 Copilot?
Possibly. Run the same 40 questions against it before deciding. If it handles the addendum, conflict and permission questions, configuring what you already pay for is the cheaper route. If it fails them, the failed questions become your requirements for anything else you consider. The broader choice between integrating, replacing and building is covered in our guide to custom software versus off-the-shelf tools.
How do we keep the system honest after launch?
Rerun the evaluation set after any change to chunking, the model, index settings or permission sync. Add every real failure to it, with the expected answer checked by an engineer. It also helps when answers say how much to trust them. Our account of making an AI system flag when it might be wrong shows one way to build that for numerical data.
Bring one package and ten questions
If you want a second opinion on a system you're evaluating or planning, bring one complete bid package with its addenda, plus ten questions your engineers actually asked on it. We would look first at how the specifications are split into passages and whether each addendum item is linked to the paragraph it replaces, because those two things decide whether a cited answer can be trusted. That is the starting point for any internal AI system we build over engineering documents.