Common Applied AI and Generative AI interview questions — LLMs, RAG, agents, embeddings and evaluation — with clear answers. Practise them in a free AI mock interview.
Practise these in a free AI mock interviewRAG retrieves relevant documents from a knowledge base (via embeddings + a vector database) and feeds them into the LLM prompt so it answers from your data. It reduces hallucination and uses fresh or private information without retraining.
Use RAG when the model needs current or proprietary facts and easy updates. Fine-tune to change style, format or task behaviour that examples can teach. Most systems try RAG first (cheaper, updatable) and fine-tune only when needed.
A numeric vector that captures the meaning of text (or images) so similar items sit close together in vector space. Embeddings power semantic search and RAG retrieval.
Ground answers with RAG, tell the model to say "I don't know" when unsure, lower temperature for factual tasks, add citations, and validate outputs with schemas/checks or an evaluator. Test on a fixed eval set before shipping.
LLMs process text as tokens — sub-word chunks, so a common word may be one token and a rare word several. Cost and context limits are measured in tokens, not words.
Designing the instructions and context you give an LLM to get reliable output — being specific, giving examples (few-shot), setting a role, requesting a format like JSON, and breaking complex tasks into steps.
An LLM that decides and takes actions using tools (search, code, APIs) in a loop to reach a goal, rather than just returning text. Agents plan, call tools, observe results and iterate.
Build a representative test set, define metrics (accuracy, relevance, format-correctness, safety), run automated evals including LLM-as-judge, review edge cases manually, and monitor in production. Never ship on a single good demo.
Use smaller/cheaper models where they suffice, cache repeated calls, trim prompt and context length, cap max tokens, batch requests, and add per-user rate limits.
A database optimised to store embeddings and find the nearest (most similar) vectors quickly. It is the retrieval backbone of RAG and semantic search.
A setting that controls randomness. A low temperature near 0 gives focused, repeatable answers — good for factual tasks. A higher temperature gives more varied, creative output.
The maximum amount of text, measured in tokens, a model can consider at once — prompt plus response. Go over it and older content is dropped, so long inputs need summarising or retrieval.
Zero-shot gives only instructions; few-shot adds a few worked examples so the model infers the pattern. Examples improve reliability on formatting and tricky edge cases.
Asking the model to reason step by step before answering. It improves accuracy on multi-step problems, though in production you often keep only the final answer and hide the reasoning.
Letting the model output a structured request to call a function or API with arguments, which your code runs and returns. It is how agents take real actions like searching or querying a database.
When untrusted input tricks the model into ignoring its instructions, for example "ignore previous instructions". Defend by separating system instructions from user data, validating inputs, limiting tool permissions, and checking outputs.
Use RAG to supply source passages, instruct the model to answer only from them and cite the source, and have it say it does not know when the passages do not contain the answer.
Using an LLM to score or compare other outputs against criteria. It scales evaluation cheaply but can be biased or inconsistent, so you validate it against human labels.
Training a base model further on your own examples to change its style or behaviour. It is worth it when prompting and RAG cannot reach the consistency you need, you have quality data, and the task is stable.
Reducing the numeric precision of a model's weights, for example from 16-bit to 4-bit, to shrink memory and speed up inference with a small accuracy trade-off. It lets larger models run on smaller hardware.
Stream the response, use a smaller or faster model where it suffices, cache, run retrieval in parallel, keep prompts short, and show partial output so the app feels responsive.
When a model learns the training data too closely, including its noise, and then performs poorly on new data. Prevent it with more data, regularisation, simpler models and proper validation.
Classification predicts a category (spam or not spam); regression predicts a continuous number (a price). The choice shapes the model, the loss function and the metrics.
Ingest and chunk the documents, embed them into a vector store, retrieve the most relevant chunks per question (RAG), prompt the LLM to answer from them with citations, add guardrails, and evaluate on real questions.
The system message sets the rules and role, user messages are the human input, and assistant messages are the model's replies. Keeping instructions in the system message and data in user messages improves control and safety.
Knowing the answers isn’t enough — say them out loud
Practise these Applied AI questions in a free AI mock interview: answer by voice, get instant feedback on your strengths and the gaps to fix.
Admissions open · free to apply
Attended a masterclass or have a friend's referral code? You get ₹15,000 off. Fill this and our team takes it from here — pay by cash or online.