Commits · dc7404cea34ef997dfe89ca94d16358e9d29c8d8 · linux / linux-davinci

15 Oct, 2008 40 commits

KVM: Handle spurious acks for PIT interrupts · dc7404ce

Avi Kivity authored Aug 17, 2008

Spurious acks can be generated, for example if the PIC is being reset.
Handle those acks gracefully rather than flooding the log with warnings.
Signed-off-by: Avi Kivity <avi@qumranet.com>

dc7404ce

KVM: fix i8259 reset irq acking · 85428ac7

Marcelo Tosatti authored Aug 14, 2008

The irq ack during pic reset has three problems:

- Ignores slave/master PIC, using gsi 0-8 for both.
- Generates an ACK even if the APIC is in control.
- Depends upon IMR being clear, which is broken if the irq was masked
at the time it was generated.

The last one causes the BIOS to hang after the first reboot of
Windows installation, since PIT interrupts stop.

[avi: fix check whether pic interrupts are seen by cpu]
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

85428ac7

KVM: Simplify exception entries by using __ASM_SIZE and _ASM_PTR · 8ceed347
Avi Kivity authored Aug 14, 2008
```
Signed-off-by: Avi Kivity <avi@qumranet.com>
```
8ceed347
KVM: VMX: Use interrupt queue for !irqchip_in_kernel · ecfc79c7
Avi Kivity authored Aug 14, 2008
```
Signed-off-by: Avi Kivity <avi@qumranet.com>
```
ecfc79c7

KVM: set debug registers after "schedulable" section · 29415c37

Marcelo Tosatti authored Aug 01, 2008

The vcpu thread can be preempted after the guest_debug_pre() callback,
resulting in invalid debug registers on the new vcpu.

Move it inside the non-preemptable section.
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

29415c37

KVM: remove unused field from the assigned dev struct · 8349b5cd

Ben-Ami Yassour authored Aug 05, 2008

Remove unused field: struct kvm_assigned_pci_dev assigned_dev
from struct: struct kvm_assigned_dev_kernel
Signed-off-by: Ben-Ami Yassour <benami@il.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

8349b5cd

KVM: VMX: Clean up magic number 0x66 in init_rmode_tss · 464d17c8
Sheng Yang authored Aug 13, 2008
```
Signed-off-by: Sheng Yang <sheng.yang@intel.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>
```
464d17c8

KVM: Reduce stack usage in kvm_pv_mmu_op() · 6ad18fba

Dave Hansen authored Aug 11, 2008

We're in a hot path.  We can't use kmalloc() because
it might impact performance.  So, we just stick the buffer that
we need into the kvm_vcpu_arch structure.  This is used very
often, so it is not really a waste.

We also have to move the buffer structure's definition to the
arch-specific x86 kvm header.
Signed-off-by: Dave Hansen <dave@linux.vnet.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

6ad18fba

KVM: Reduce stack usage in kvm_arch_vcpu_ioctl() · b772ff36

Dave Hansen authored Aug 11, 2008

[sheng: fix KVM_GET_LAPIC using wrong size]
Signed-off-by: Dave Hansen <dave@linux.vnet.ibm.com>
Signed-off-by: Sheng Yang <sheng.yang@intel.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

b772ff36

KVM: Reduce stack usage in kvm_vcpu_ioctl() · fa3795a7

Dave Hansen authored Aug 11, 2008

Signed-off-by: Dave Hansen <dave@linux.vnet.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

fa3795a7

KVM: Reduce kvm stack usage in kvm_arch_vm_ioctl() · f0d66275

Dave Hansen authored Aug 11, 2008

On my machine with gcc 3.4, kvm uses ~2k of stack in a few
select functions.  This is mostly because gcc fails to
notice that the different case: statements could have their
stack usage combined.  It overflows very nicely if interrupts
happen during one of these large uses.

This patch uses two methods for reducing stack usage.
1. dynamically allocate large objects instead of putting
   on the stack.
2. Use a union{} member for all of the case variables. This
   tricks gcc into combining them all into a single stack
   allocation. (There's also a comment on this)
Signed-off-by: Dave Hansen <dave@linux.vnet.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

f0d66275

KVM: pci device assignment · 4d5c5d0f

Ben-Ami Yassour authored Jul 28, 2008

Based on a patch from: Amit Shah <amit.shah@qumranet.com>

This patch adds support for handling PCI devices that are assigned to
the guest.

The device to be assigned to the guest is registered in the host kernel
and interrupt delivery is handled.  If a device is already assigned, or
the device driver for it is still loaded on the host, the device
assignment is failed by conveying a -EBUSY reply to the userspace.

Devices that share their interrupt line are not supported at the moment.

By itself, this patch will not make devices work within the guest.
The VT-d extension is required to enable the device to perform DMA.
Another alternative is PVDMA.
Signed-off-by: Amit Shah <amit.shah@qumranet.com>
Signed-off-by: Ben-Ami Yassour <benami@il.ibm.com>
Signed-off-by: Weidong Han <weidong.han@intel.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

4d5c5d0f

KVM: direct mmio pfn check · cbff90a7

Ben-Ami Yassour authored Jul 28, 2008

Userspace may specify memory slots that are backed by mmio pages rather than
normal RAM.  In some cases it is not enough to identify these mmio pages
by pfn_valid().  This patch adds checking the PageReserved as well.
Signed-off-by: Ben-Ami Yassour <benami@il.ibm.com>
Signed-off-by: Muli Ben-Yehuda <muli@il.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

cbff90a7

x86: KVM guest: use paravirt function to calculate cpu khz · 0293615f

Glauber Costa authored Jul 28, 2008

We're currently facing timing problems in guests that do
calibration under heavy load, and then the load vanishes.
This means we'll have a much lower lpj than we actually should,
and delays end up taking less time than they should, which is a
nasty bug.

Solution is to pass on the lpj value from host to guest, and have it
preset.
Signed-off-by: Glauber Costa <gcosta@redhat.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

0293615f

x86: paravirt: factor out cpu_khz to common code · 3807f345

Glauber Costa authored Jul 28, 2008

KVM intends to use paravirt code to calibrate khz. Xen
current code will do just fine. So as a first step, factor out
code to pvclock.c.
Signed-off-by: Glauber Costa <gcosta@redhat.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

3807f345

KVM: PIT: fix injection logic and count · 3cf57fed

Marcelo Tosatti authored Jul 26, 2008

The PIT injection logic is problematic under the following cases:

1) If there is a higher priority vector to be delivered by the time
kvm_pit_timer_intr_post is invoked ps->inject_pending won't be set.
This opens the possibility for missing many PIT event injections (say if
guest executes hlt at this point).

2) ps->inject_pending is racy with more than two vcpus. Since there's no locking
around read/dec of pt->pending, two vcpu's can inject two interrupts for a single
pt->pending count.

Fix 1 by using an irq ack notifier: only reinject when the previous irq
has been acked. Fix 2 with appropriate locking around manipulation of
pending count and irq_ack by the injection / ack paths.
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

3cf57fed

KVM: irq ack notification · f5244726

Marcelo Tosatti authored Jul 26, 2008

Based on a patch from: Ben-Ami Yassour <benami@il.ibm.com>
which was based on a patch from: Amit Shah <amit.shah@qumranet.com>

Notify IRQ acking on PIC/APIC emulation. The previous patch missed two things:

- Edge triggered interrupts on IOAPIC
- PIC reset with IRR/ISR set should be equivalent to ack (LAPIC probably
needs something similar).
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>
CC: Amit Shah <amit.shah@qumranet.com>
CC: Ben-Ami Yassour <benami@il.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

f5244726

KVM: Add irq ack notifier list · 564f1537

Avi Kivity authored Jul 26, 2008

This can be used by kvm subsystems that are interested in when
interrupts are acked, for example time drift compensation.
Signed-off-by: Avi Kivity <avi@qumranet.com>

564f1537

KVM: powerpc: Map guest userspace with TID=0 mappings · 49dd2c49

Hollis Blanchard authored Jul 25, 2008

When we use TID=N userspace mappings, we must ensure that kernel mappings have
been destroyed when entering userspace. Using TID=1/TID=0 for kernel/user
mappings and running userspace with PID=0 means that userspace can't access the
kernel mappings, but the kernel can directly access userspace.

The net is that we don't need to flush the TLB on privilege switches, but we do
on guest context switches (which are far more infrequent). Guest boot time
performance improvement: about 30%.
Signed-off-by: Hollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

49dd2c49

KVM: ppc: Write only modified shadow entries into the TLB on exit · 83aae4a8

Hollis Blanchard authored Jul 25, 2008

Track which TLB entries need to be written, instead of overwriting everything
below the high water mark. Typically only a single guest TLB entry will be
modified in a single exit.

Guest boot time performance improvement: about 15%.
Signed-off-by: Hollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

83aae4a8

KVM: ppc: Stop saving host TLB state · 20754c24

Hollis Blanchard authored Jul 25, 2008

We're saving the host TLB state to memory on every exit, but never using it.
Originally I had thought that we'd want to restore host TLB for heavyweight
exits, but that could actually hurt when context switching to an unrelated host
process (i.e. not qemu).

Since this decreases the performance penalty of all exits, this patch improves
guest boot time by about 15%.
Signed-off-by: Hollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

20754c24

KVM: ppc: guest breakpoint support · 6a0ab738

Hollis Blanchard authored Jul 25, 2008

Allow host userspace to program hardware debug registers to set breakpoints
inside guests.
Signed-off-by: Jerone Young <jyoung5@us.ibm.com>
Signed-off-by: Hollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

6a0ab738

KVM: Ignore DEBUGCTL MSRs with no effect · b5e2fec0

Alexander Graf authored Jul 22, 2008

Netware writes to DEBUGCTL and reads from the DEBUGCTL and LAST*IP MSRs
without further checks and is really confused to receive a #GP during that.
To make it happy we should just make them stubs, which is exactly what SVM
already does.

Writes to DEBUGCTL that are vendor-specific are resembled to behave as if the
virtual CPU does not know them.
Signed-off-by: Alexander Graf <agraf@suse.de>
Signed-off-by: Avi Kivity <avi@qumranet.com>

b5e2fec0

KVM: VMX: Avoid vmwrite(HOST_RSP) when possible · 313dbd49

Avi Kivity authored Jul 17, 2008

Usually HOST_RSP retains its value across guest entries.  Take advantage
of this and avoid a vmwrite() when this is so.
Signed-off-by: Avi Kivity <avi@qumranet.com>

313dbd49

KVM: ppc: trace powerpc instruction emulation · 3b4bd796

Christian Ehrhardt authored Jul 14, 2008

This patch adds a trace point for the instruction emulation on embedded powerpc
utilizing the KVM_TRACE interface.
Signed-off-by: Christian Ehrhardt <ehrhardt@linux.vnet.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

3b4bd796

KVM: ppc: adds trace points for ppc tlb activity · 31711f22

Jerone Young authored Jul 14, 2008

This patch adds trace points to track powerpc TLB activities using the
KVM_TRACE infrastructure.
Signed-off-by: Jerone Young <jyoung5@us.ibm.com>
Signed-off-by: Christian Ehrhardt <ehrhardt@linux.vnet.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

31711f22

KVM: ppc: enable KVM_TRACE building for powerpc · 12f67556

Jerone Young authored Jul 14, 2008

This patch enables KVM_TRACE to build for PowerPC arch. This means just
adding sections to Kconfig and Makefile.
Signed-off-by: Jerone Young <jyoung5@us.ibm.com>
Signed-off-by: Christian Ehrhardt <ehrhardt@linux.vnet.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

12f67556

KVM: kvmtrace: replace get_cycles with ktime_get v3 · 3f7f95c6

Christian Ehrhardt authored Jul 14, 2008

The current kvmtrace code uses get_cycles() while the interpretation would be
easier using using nanoseconds. ktime_get() should give at least the same
accuracy as get_cycles on all architectures (even better on 32bit archs) but
at a better unit (e.g. comparable between hosts with different frequencies.

[avi: avoid ktime_t in public header]
Signed-off-by: Christian Ehrhardt <ehrhardt@linux.vnet.ibm.com>
Acked-by: Christian Borntraeger <borntraeger@de.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

3f7f95c6

KVM: kvmtrace: Remove use of bit fields in kvm trace structure · e32c8f2c

Christian Ehrhardt authored Jul 14, 2008

This patch fixes kvmtrace use on big endian systems. When using bit fields the
compiler will lay data out in the wrong order expected when laid down into a
file.
This fixes it by using one variable instead of using bit fields.
Signed-off-by: Jerone Young <jyoung5@us.ibm.com>
Signed-off-by: Christian Ehrhardt <ehrhardt@linux.vnet.ibm.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

e32c8f2c

KVM: SVM: Unify register save/restore across 32 and 64 bit hosts · 80e31d4f
Avi Kivity authored Jul 14, 2008
```
Signed-off-by: Avi Kivity <avi@qumranet.com>
```
80e31d4f
KVM: VMX: Unify register save/restore across 32 and 64 bit hosts · c801949d
Avi Kivity authored Jul 14, 2008
```
Signed-off-by: Avi Kivity <avi@qumranet.com>
```
c801949d

KVM: VMX: Reinject real mode exception · 77ab6db0

Jan Kiszka authored Jul 14, 2008

As we execute real mode guests in VM86 mode, exception have to be
reinjected appropriately when the guest triggered them. For this purpose
the patch adopts the real-mode injection pattern used in vmx_inject_irq
to vmx_queue_exception, additionally taking care that the IP is set
correctly for #BP exceptions. Furthermore it extends
handle_rmode_exception to reinject all those exceptions that can be
raised in real mode.

This fixes the execution of himem.exe from FreeDOS and also makes its
debug.com work properly.

Note that guest debugging in real mode is broken now. This has to be
fixed by the scheduled debugging infrastructure rework (will be done
once base patches for QEMU have been accepted).
Signed-off-by: Jan Kiszka <jan.kiszka@web.de>
Signed-off-by: Avi Kivity <avi@qumranet.com>

77ab6db0

KVM: Consolidate XX_VECTOR defines · 19bd8afd

Jan Kiszka authored Jul 13, 2008

Signed-off-by: Jan Kiszka <jan.kiszka@web.de>
Signed-off-by: Avi Kivity <avi@qumranet.com>

19bd8afd

KVM: Consolidate PIC isr clearing into a function · 7edd0ce0
Avi Kivity authored Jul 07, 2008
```
Signed-off-by: Avi Kivity <avi@qumranet.com>
```
7edd0ce0

KVM: VMX: Remove redundant check in handle_rmode_exception · 60bd83a1

Mohammed Gamal authored Jul 12, 2008

Since checking for vcpu->arch.rmode.active is already done whenever we
call handle_rmode_exception(), checking it inside the function is redundant.
Signed-off-by: Mohammed Gamal <m.gamal005@gmail.com>
Signed-off-by: Avi Kivity <avi@qumranet.com>

60bd83a1

KVM: VMX: Move interrupt post-processing to vmx_complete_interrupts() · f7d9238f

Avi Kivity authored Jul 03, 2008

Instead of looking at failed injections in the vm entry path, move
processing to the exit path in vmx_complete_interrupts(). This simplifes
the logic and removes any state that is hidden in vmx registers.
Signed-off-by: Avi Kivity <avi@qumranet.com>

f7d9238f

KVM: Add a pending interrupt queue · 937a7eae

Avi Kivity authored Jul 03, 2008

Similar to the exception queue, this hold interrupts that have been
accepted by the virtual processor core but not yet injected.

Not yet used.
Signed-off-by: Avi Kivity <avi@qumranet.com>

937a7eae

KVM: VMX: Fix pending exception processing · 35920a35

Avi Kivity authored Jul 03, 2008

The vmx code assumes that IDT-Vectoring can only be set when an exception
is injected due to the exception in question.  That's not true, however:
if the exception is injected correctly, and later another exception occurs
but its delivery is blocked due to a fault, then we will incorrectly assume
the first exception was not delivered.

Fix by unconditionally dequeuing the pending exception, and requeuing it
(or the second exception) if we see it in the IDT-Vectoring field.
Signed-off-by: Avi Kivity <avi@qumranet.com>

35920a35

KVM: Clear exception queue before emulating an instruction · 26eef70c

Avi Kivity authored Jul 03, 2008

If we're emulating an instruction, either it will succeed, in which case
any previously queued exception will be spurious, or we will requeue the
same exception.
Signed-off-by: Avi Kivity <avi@qumranet.com>

26eef70c

KVM: VMX: Move nmi injection failure processing to vm exit path · 668f612f

Avi Kivity authored Jul 02, 2008

Instead of processing nmi injection failure in the vm entry path, move
it to the vm exit path (vm_complete_interrupts()).  This separates nmi
injection from nmi post-processing, and moves the nmi state from the VT
state into vcpu state (new variable nmi_injected specifying an injection
in progress).
Signed-off-by: Avi Kivity <avi@qumranet.com>

668f612f