When interviewers ask you to compare fine‑tuning and prompting, they want to see that you understand both the technical underpinnings and the practical trade‑offs. Below is a concise framework you can use, plus a ready‑to‑speak answer.
One‑sentence definitions
- Fine‑tuning: Updating a pre‑trained model’s weights on a curated dataset so it behaves better for a specific domain or task.
- Prompting: Providing a natural‑language instruction (or few‑shot examples) that guides the same model to produce the desired output without changing its weights.
How each mechanism works
Fine‑tuning
- Start with a base model (e.g., a 7B transformer trained on internet text).
- Collect labeled data that reflects the target use case.
- Run gradient descent to adjust the model parameters, typically for a few epochs.
- Validate on a held‑out set to ensure the model has learned the new behavior.
Prompting
- Write a prompt that frames the task, optionally including examples (few‑shot) or chain‑of‑thought reasoning.
- Send the prompt to the unchanged model via an API call.
- Parse the response; you may need post‑processing or a second prompt to refine it.
Trade‑offs
| Aspect | Fine‑tuning | Prompting |
|---|---|---|
| Performance on niche tasks | Typically higher, because the model sees many examples of the target pattern. | Variable; depends on prompt quality and model size. |
| Resource cost | Requires GPU time, storage for datasets, and engineering effort. | Minimal compute beyond the inference call. |
| Maintenance | Needs versioning, re‑training when data drifts. | Only the prompt text changes; easier to iterate. |
| Speed to ship | Weeks to months for data prep, training, and testing. | Hours or less; you can prototype by editing text. |
| Risk of over‑fitting | Possible if data is small or not diverse. | Not applicable; model stays general. |
| Explainability | You can inspect fine‑tuned weights, but they are still opaque. | Prompt logic is visible and can be reviewed. |
Concrete example
Imagine a company that processes medical insurance claims.
- Fine‑tuning: They gather 10 k annotated claim forms and fine‑tune a 13B model to extract fields like patient ID, diagnosis code, and amount. After training, the model achieves >90 % field‑level accuracy on their internal test set.
- Prompting: The same team writes a prompt: "Extract the patient ID, diagnosis code, and claim amount from the following insurance claim.\n\n[Claim text]" They send it to the base model and get results that are correct about 70 % of the time, but they can improve it by adding a few examples in‑prompt. The fine‑tuned model is more reliable for high‑volume batch processing, while prompting is handy for ad‑hoc queries or rapid prototyping.
Typical interview questions
- Why would you choose prompting over fine‑tuning for a new feature? Answer: Prompting is faster, cheaper, and lets you test ideas without committing resources. It’s ideal when the use case is short‑lived or the data is scarce.
- What are the main cost considerations? Answer: Fine‑tuning incurs GPU hours, data labeling, and ongoing model management; prompting only costs inference time and prompt engineering effort.
- How do you guard against prompt injection attacks? Answer: Validate user‑supplied content, use sandboxed contexts, and avoid concatenating raw input directly into the prompt.
- When does fine‑tuning become necessary? Answer: When the baseline model consistently fails on domain‑specific language, or when latency and error rates matter for production.
- Can you combine both? Answer: Yes—fine‑tune on a core dataset to get a strong baseline, then use prompting for edge cases or to inject up‑to‑date policies.
60‑second spoken answer (ready for the interview)
"Fine‑tuning and prompting are two ways to adapt a large language model to a task. Fine‑tuning modifies the model’s weights by training on a domain‑specific dataset; it gives higher accuracy on niche problems but requires data, compute, and ongoing maintenance. Prompting, on the other hand, leaves the model unchanged and steers it with a carefully crafted instruction or few‑shot examples. It’s cheap and fast, but performance can be variable. For example, a health‑insurer might fine‑tune a model on 10 k annotated claim forms to achieve >90 % extraction accuracy, while a prototype team could just prompt the base model to extract fields and get about 70 % accuracy. In practice, you pick prompting for quick experiments or low‑risk features, and you move to fine‑tuning when you need consistent, high‑throughput results. I’d also watch for prompt injection risks and be ready to fall back to a fine‑tuned model if the baseline keeps failing."
How to practice this
- Write your own prompt for a simple task (e.g., summarizing a news article) and measure the output quality.
- Create a tiny fine‑tuning dataset (20–30 examples) and run a quick LoRA or adapter training on a public model; compare the results to the prompt.
- Simulate an interview: use Call Assistant to record yourself delivering the 60‑second answer, then replay the transcript to check clarity and timing.
FAQ
- Q: Is fine‑tuning always better than prompting? A: Not necessarily. Fine‑tuning gives higher performance on specialized data, but it costs more time and resources. Prompting can be sufficient for many tasks, especially when you need speed and flexibility.
- Q: What is a "few‑shot" prompt? A: It includes a small number of input‑output pairs in the prompt to illustrate the pattern you want the model to follow.
- Q: How do I know when a prompt is too long? A: Most APIs have token limits (e.g., 8 k tokens). If you exceed it, the model will truncate, so keep prompts concise and use external context storage when needed.
- Q: Can I fine‑tune a model without a GPU? A: You can use parameter‑efficient methods like LoRA that run on modest hardware, but training still benefits from a GPU for speed.
Frequently asked questions
Is fine‑tuning always better than prompting?
Not necessarily. Fine‑tuning gives higher performance on specialized data, but it costs more time and resources. Prompting can be sufficient for many tasks, especially when you need speed and flexibility.
What is a few‑shot prompt?
A few‑shot prompt includes a handful of input‑output examples inside the prompt to illustrate the desired pattern, helping the model infer the task without weight updates.
How do I know when a prompt is too long?
Check the token limit of the API you’re using (commonly 8 k tokens). If the prompt plus expected response exceeds that, the model will truncate, so keep prompts concise and move large context to external storage.
Can I fine‑tune a model without a GPU?
Parameter‑efficient techniques like LoRA can run on modest CPUs, but training will be slower. A GPU still provides a noticeable speed advantage for most fine‑tuning workloads.
#concept#fine-tuning vs prompting#interview#AI#ML